How to Use Gemini Agentic Video Understanding
Learn how to use Gemini's agentic video understanding to search long videos, find precise moments, inspect fast actions, and reduce unnecessary token usage.
On this page
Long videos are expensive to analyze when an AI model has to inspect them at a fixed frame rate from beginning to end. Google's agentic video understanding changes that workflow by letting Gemini decide which parts of a video deserve closer inspection. The feature, released for the Gemini API on September 1, 2026, can dynamically request frames, audio, or transcripts instead of treating the entire video as a uniform stream.
What agentic video understanding changes
Traditional video analysis commonly samples a video at a fixed rate, such as one frame per second. That approach is predictable, but it can waste tokens on unimportant sections while missing a short event that happens between sampled frames. Agentic video understanding gives Gemini an internal video tool and lets the model decide what section to inspect, how closely to inspect it, and whether visual frames, audio, or a transcript is most useful.
Google says the feature can reduce token consumption by up to 88%, reduce analysis costs by up to 66%, and improve accuracy by up to 7% on its tested video-analysis benchmarks. Those are Google's benchmark results, not a guarantee for every video or application, so your own workload should still be measured before you assume the same savings.
Set up the Gemini Python SDK
The easiest way to experiment with the feature is through Google's Python SDK. Install the current GenAI package and configure your Gemini API key through an environment variable so the credential does not become part of your application source.
pip install google-genaiThen create the client in Python. The SDK reads the API credential from the environment when GEMINI_API_KEY is configured.
from google import genai
client = genai.Client()This separates authentication from the code that performs the video analysis. For production applications, keep the API key in your server-side secret-management system rather than shipping it inside browser or mobile application code.
Analyze a YouTube video with agentic processing
The current Gemini example uses the Interactions API. You provide a video input, set its processing mode to agentic, and add the question you want the model to answer. Gemini then determines which portions of the video it needs to inspect.
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.7-flash",
input=[
{
"type": "video",
"uri": "[https://youtu.be/VIDEO_ID](https://youtu.be/VIDEO_ID)",
"processing": "agentic"
},
{
"type": "text",
"text": "What are the three most important announcements in this keynote?"
},
],
)
print(interaction.output_text)The important setting is "processing": "agentic". Without that setting, you are not explicitly asking the API to use the new agentic video-processing path. Google's launch documentation shows this configuration with Gemini 3.7 Flash and a YouTube video URL. ([blog.google][1])
Ask for a specific moment instead of a general summary
Agentic video understanding becomes more useful when the question requires the model to locate something precise. A generic summary can often be produced from broad sampling, while a question about a brief event requires the system to find the relevant portion of the timeline.
interaction = client.interactions.create(
model="gemini-3.7-flash",
input=[
{
"type": "video",
"uri": "[https://youtu.be/VIDEO_ID](https://youtu.be/VIDEO_ID)",
"processing": "agentic"
},
{
"type": "text",
"text": (
"Find every moment where the presenter changes "
"the slide and mention the timestamp and topic."
)
},
],
)
print(interaction.output_text)The model can dynamically inspect the timeline rather than requiring your application to split the entire video into thousands of manually selected frames. Google specifically lists sub-second moment retrieval as one of the intended uses of the technology. ([blog.google][1])
Use agentic processing for fast actions
Fixed-rate sampling can struggle with rapid movements. If a person performs an action in a fraction of a second, a one-frame-per-second sampling strategy may never capture the critical state clearly. Agentic processing can resample an interesting section at a higher frame rate when the initial inspection suggests that more detail is needed.
This makes the approach useful for applications such as sports analysis, automated editing, industrial inspection, and counting repeated physical actions. Google demonstrates the approach with fast-paced movement counting, where the model can inspect the same section at different frame rates as needed. ([blog.google][1])
Use transcripts and audio when the picture is not enough
A video contains more than images. A speaker may explain an event verbally while the important visual detail appears elsewhere, so an effective video-analysis system sometimes needs both signals. Agentic video understanding can request frames, audio, or transcripts based on what the question requires.
For example, instead of asking the model to describe every frame in a two-hour lecture, ask it to find where the speaker explains a particular concept. The system can use the transcript to locate likely sections and then inspect the corresponding visual material when the answer depends on what appeared on screen.
Analyze long videos without processing everything equally
The biggest practical difference appears with long-form material. Google says the system is designed for content ranging from 10-minute how-to videos and 90-minute lectures to multi-hour recordings. Static processing forces developers to balance sampling quality against token usage, while the agentic approach can concentrate processing on sections relevant to the question. ([blog.google][1])
That does not mean a long video becomes free to analyze. The model still has to inspect content, and complicated questions can require substantial processing. The advantage is that the system can avoid spending equal effort on every second of the recording when only a small portion contains the answer.
Build an anomaly-detection workflow
Anomaly detection is another natural fit because unusual events often occupy only a small portion of a recording. Rather than processing every moment at maximum detail, an application can ask Gemini to locate suspicious sections and inspect those sections more closely.
question = """
Review this recording for unusual visual events.
For each possible anomaly, report:
1. approximate timestamp,
2. what changed,
3. why the event appears unusual.
"""This type of workflow should still be treated as an AI-assisted detection system rather than an automatic proof that an event is abnormal. For security, safety, or industrial applications, detected events should be validated against the requirements of the specific environment.
Keep Google's benchmark claims in context
Google reports up to 88% lower token consumption, up to 66% lower cost, and up to 7% higher accuracy across its tested benchmarks. The company also highlights long-form video as an area where the efficiency difference is particularly useful. ([blog.google][1])
Those figures describe Google's own evaluation. An independent developer discussion has reported mixed results on a small synthetic benchmark, including cases where static processing used fewer tokens and cost less. That test was explicitly described by its author as exploratory, with only a small number of videos, so it should not be treated as a universal counter-result. ([Google AI Developers Forum][2])
The practical lesson is to benchmark your own workload. If your application mostly asks broad questions that can be answered from coarse sampling, the benefit may be smaller. If it repeatedly searches long recordings for brief events, the dynamic inspection model may provide a more meaningful difference.
Choose the right Gemini model
Google launched agentic video understanding across Gemini 3.7 Flash, Gemini 3.6 Flash, and Gemini 3.5 Flash-Lite. The capability is available through the Gemini API and can be used with video uploads as well as YouTube videos. ([blog.google][1])
For a first prototype, use one supported model and measure the results before changing models. Keep the video, question, and evaluation criteria consistent when comparing configurations. That makes it easier to determine whether a change actually improved your application instead of simply producing a different answer.
Test the workflow before putting it into production
Create a small evaluation set containing videos where you already know the correct timestamps, counts, or events. Include easy cases and deliberately difficult ones, such as short events, rapid movement, speech over visual changes, and long periods containing nothing relevant.
Measure more than answer quality. Record token usage, processing time, cost, missed events, false detections, and whether the model identifies the correct part of the video. A system that saves tokens but consistently misses the event your users care about is not solving the underlying problem.
What to build first with agentic video
A good first project is a narrow video-search tool. Let a user upload a recording or provide a supported video source, ask a specific question, and return the relevant moments with a short explanation. Once that works reliably, you can add features such as event counting, anomaly detection, automated editing assistance, or searchable video archives.
The key change is that your application no longer has to decide in advance exactly which frames Gemini should inspect. You describe the video task, enable agentic processing, and let the model navigate the timeline according to the question. For long or event-heavy recordings, that turns video analysis from a fixed sampling problem into a targeted search problem.
Written by


