Introducing agentic video understanding with Gemini
| Source: Google DeepMind Blog
Tags: Gemini, Google DeepMind, video AI, multimodal, Gemini 3.7 Flash, cost reduction
Google DeepMind launched agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, replacing fixed-rate frame ingestion with dynamic scanning across video, audio, and transcripts — cutting token usage by up to 88%, costs by up to 66%, and boosting accuracy by up to 7% on benchmarks.
Details
Google DeepMind has shipped agentic video understanding across three Gemini models (3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite), available now via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. The feature works with both video file uploads and YouTube URLs, activated by setting the API configuration to 'agentic'.\n\nThe core shift is from static frame-rate ingestion (default 1 FPS) to an agentic approach where the model dynamically searches, scans, and inspects target video segments across visual frames, audio, and transcripts. This parallels the earlier 'agentic vision' feature — which combines code execution with image understanding — but applied to video.\n\nBenchmarks show genuine Pareto improvements: up to 88% fewer tokens, up to 66% lower cost, and up to 7% better accuracy simultaneously. The gains are most pronounced on long-form content (10-minute how-to guides to multi-hour recordings), where static processing forced developers to choose between high token spend or dropping critical details. Gemini 3.7 Flash sits at the accuracy-to-cost Pareto frontier among tested models.\n\nNew capabilities unlocked include sub-second moment retrieval, more accurate anomaly detection, and precise object counting. For teams running video pipelines at scale, the efficiency gains are immediate and substantial — no model change required, just a configuration switch.