Introducing Agentic Video Understanding with Gemini: Up to 88% Fewer Tokens, 66% Lower Cost
Google DeepMind has launched agentic video understanding for Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. Instead of ingesting video at a fixed frame rate, the model now searches, scans and inspects relevant segments across frames, audio and transcripts in an agentic loop. Google reports up to 88% fewer tokens, up to 66% lower cost and up to 7% higher accuracy on standard video benchmarks. Gains are largest on long video. Developers enable it by setting the API configuration to 'agentic' in Google AI Studio or the Gemini Enterprise Agent Platform.
Google DeepMind has launched agentic video understanding for Gemini, and made it available on three models: Gemini 3.7 Flash, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. The announcement, written by Senior Product Manager Rohan Doshi and Research Director Mario Lučić, carries three headline numbers. On standard video analysis benchmarks, the feature cuts token consumption by up to 88%, reduces analysis cost by up to 66%, and raises accuracy by up to 7%. All three figures are ceilings, not averages. Real results will depend on video length and task type. To see why this matters, start with how video is processed today. In the 'static' approach, the model ingests the video at a fixed frames-per-second rate. The default is 1 FPS, and developers can change it through the API. That is simple, but it has two costs. First, token use grows with video length, so a 90-minute lecture or a multi-hour recording is expensive. Second, developers must choose between paying high token bills and using shortcuts such as frame dropping, truncation or summarization, which can lose critical details. Google says this trade-off is sharpest on long-form video, from 10-minute how-to guides to 90-minute lectures and multi-hour recordings. That is also where the new feature helps most.
The mechanism is a change of role for the model. Instead of receiving every frame passively, Gemini pairs its core reasoning with native video tools. It can search, scan and inspect target segments across three kinds of signal: visual frames, audio and transcripts. The model decides what to watch, at what speed, and through which modality. It then fetches only the moments and signals it needs. Google describes this as an agentic loop. Gemini invokes an internal tool to load the relevant part of the video file, looks at the result, and decides what to do next. Developers could build something like this by hand before, for example with their own slicing, transcription and retrieval code. Now the loop runs inside the model, which reduces development overhead. The idea follows agentic vision, an earlier Gemini feature that combines code execution with the models' native image understanding. The new feature applies the same pattern to video. Google lists new capabilities that come with it: sub-second moment retrieval, more accurate anomaly detection, and precise counting. These tasks share one trait. The answer often sits in a few short moments, and uniform sampling at 1 FPS can miss them. A model that can go back, slow down and look again is better placed to find them. On performance, Google says the gains span all three supported models. It singles out Gemini 3.7 Flash with agentic understanding as the best overall quality, and as the best mix of quality and cost. In Google's chart it sits on the accuracy-to-cost Pareto frontier among the tested models. In plain terms, within the tested set, no other configuration gave higher accuracy at lower cost. A caution applies. The results come from standard video analysis benchmarks chosen by Google. The source text we reviewed does not name each benchmark or give per-benchmark tables. Teams should therefore test on their own videos before they trust the headline figures.
Access is simple. The feature works with video uploads and YouTube videos, through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. According to Google's summary, a developer turns it on by setting the API configuration to 'agentic'. That is a small switch, and existing video pipelines should not need a rewrite. For developers and enterprises, three effects stand out. The first is cost structure. Long-video analysis has often been hard to scale because of token bills. Examples include reviewing security footage, summarizing meetings and courses, searching media archives, and sampling for compliance. A cost reduction of up to 66% can make some of these cases viable. The second is engineering effort. Frame selection, segmentation, transcript alignment and retrieval logic that teams built themselves can move into the model's own tool loop. The third is that quality and cost no longer have to trade off. Accuracy rises while cost falls, which is attractive for production use.
There are open questions. An agentic loop means several model steps per request. Google's source text gives no latency data, so teams must measure response time and run-to-run consistency. Because the model decides where to look, it may skip a key segment in rare cases. For high-stakes work, such as safety monitoring or medical imaging, human review and spot checks remain necessary. Google has also not said whether more models or input sources will be supported. Looking ahead, the release shows where multimodal competition is heading. The contest is moving from how much video a model can hold to how intelligently it looks. A larger context window does not make full frame-by-frame input cheap. Letting a model browse, locate and then inspect closely is nearer to how people work with footage. We expect other model vendors to ship similar active-retrieval features. Developer attention will likely shift from preprocessing pipelines toward evaluation and cost monitoring.
A practical plan for teams: pick a representative set of videos, split it into short and long clips, and run each with static and agentic settings. Record tokens, total cost and accuracy per task. Google says gains are strongest on long video, so expect smaller gains on short clips. Decide on migration only after that comparison.