Skip to content

Gemini’s New Video AI Stops Watching Every Second

Google’s new agentic video understanding lets Gemini decide which parts of a video to inspect, cutting token use and costs while improving results on its benchmarks.

Gemini’s New Video AI Stops Watching Every Second

On this page

Google has changed how Gemini analyzes video. Instead of feeding a model a fixed stream of frames from beginning to end, its new agentic video understanding mode lets Gemini decide which parts of a video deserve closer inspection, when to change the sampling rate, and whether it needs visual frames, audio, or a transcript. Google says the approach can cut input-token use by up to 88%, reduce analysis costs by up to 66%, and improve accuracy by up to 7% on its tested video benchmarks. 

Gemini now decides what parts of a video matter

The change is easier to understand by looking at the old workflow. Gemini's standard video processing samples footage at a fixed rate, with 1 frame per second as the default, then places those frames into the model's context for analysis. That works reasonably well for short clips, but a long lecture, meeting, tutorial, or surveillance recording can contain thousands of mostly irrelevant frames. Agentic processing changes the job from “watch everything at the same rate” to “find the evidence needed to answer this question.” 

In practice, Gemini can explore the video timeline, request a particular segment, inspect frames at a different rate, and retrieve audio or transcript information when those signals are more useful. A question about something that happened for less than a second can therefore trigger much more detailed inspection around that moment without forcing the entire recording through the same high-resolution process. Google describes this as a goal-directed approach to video analysis rather than simple frame sampling. 

The biggest change is efficiency, not simply better vision

Google's headline number is an up to 88% reduction in token consumption for long-form video. Tokens are the units an AI model processes as input, so using fewer of them generally means less computation and lower usage costs. Google also reports up to 66% lower analysis costs and up to 7% higher quality across its standard video-analysis testing. These are Google's benchmark results, not an independently reproduced industry benchmark, so they should be treated as vendor-reported performance rather than a universal guarantee. 

The savings matter most when the video is long and the question is narrow. Consider a 90-minute lecture where someone wants to know when the presenter first introduces a particular diagram. A fixed one-frame-per-second process spends resources representing the entire lecture, even though most of it is irrelevant to that question. An agentic system can search the timeline first, identify promising sections, and then inspect those sections more carefully. The result is not simply fewer frames; it is a different allocation of the model's attention.

It can also look more closely when the video demands it

Reducing the amount of video inspected sounds like it should hurt accuracy, but fixed sampling has its own weakness: important events can happen between sampled frames. Google gives fast-moving actions and precise counting as examples where agentic processing can change the sampling strategy. Gemini can revisit a section at a higher frame rate when the question requires fine-grained visual information instead of treating the entire recording uniformly. 

That makes the technology particularly interesting for tasks such as finding a specific moment in a long recording, detecting an unusual event, counting objects or actions, and answering questions that depend on both what was said and what appeared on screen. The model is effectively choosing between different forms of evidence instead of assuming that every second of video deserves identical treatment.

The feature is already available to developers

Agentic video understanding launched on September 1, 2026, for Gemini 3.7 Flash, Gemini 3.6 Flash, and Gemini 3.5 Flash-Lite through the Gemini application programming interface (API), Google's interface for sending data to its models from software. It is available for uploaded videos and YouTube videos through Google AI Studio and the Gemini Enterprise Agent Platform. Google says there is no separate feature fee; standard Gemini API token pricing applies. 

Developers explicitly select agentic processing for a video request. Google's documentation shows the mode being set to "agentic", after which the model can dynamically request video segments, transcripts, and other information during processing. The API also exposes processing steps, allowing developers to see that the model actually made video-processing requests rather than silently treating the file as a conventional static input.

Static processing still has a place

Google does not recommend replacing every video workflow with agentic processing. Its documentation says static processing can remain useful for short clips, latency-sensitive workloads, and situations where consistent frame-level coverage across the entire video matters. Agentic processing is better suited to long-form recordings or questions aimed at particular moments.

That distinction is important for developers designing production systems. A five-minute clip that needs a complete frame-by-frame inspection may not benefit from an agent deciding where to look. A two-hour recording containing a handful of relevant moments is a very different problem. Choosing the processing mode based on the task, rather than assuming the newest option is always better, should produce more predictable costs and results.

Google is pushing the same approach beyond the API

The developer release is only the first stage of the rollout. Google says the efficiency and quality improvements will come to users of the Gemini app across its Flash and Flash-Lite models, while agentic video understanding is also planned for YouTube's “Ask YouTube” feature in the coming months. That would move the technology from a developer-facing optimization into a consumer feature where people can ask questions about videos without manually searching through their timelines.

The more significant shift is therefore architectural. Video AI does not necessarily need to consume every part of a recording equally. By giving the model tools to search, inspect, revisit, and change how it samples footage, Google is treating a video less like a giant collection of images and more like information that an AI system can actively investigate. The next question is whether these gains hold consistently outside Google's own benchmarks and across the much wider range of videos developers encounter in production.

M

Written by

M Umar Farooq

I’m curious about new technology and the ideas that are changing how we use digital products and services. I enjoy exploring emerging technologies, useful tools, new features, and clever solutions to everyday technology problems. I especially like finding simple fixes and practical tricks that can save people time and frustration.

25 posts published

All posts by this author

0 Comments

No comments yet. Be the first to share your thoughts.

Join the conversation

Log in or create a free account to leave a comment. You can edit or delete your own comments any time.