Google has added agentic video understanding to Gemini. Instead of scanning a video from start to finish at a fixed interval, the model now goes and fetches only the moments it needs to answer the question. On standard video analysis benchmarks, token consumption drops by up to 88%, analysis costs fall by up to 66%, and accuracy improves by up to 7%[1].

Dropping the one-frame-per-second approach

Video processing used to be static. The model ingested a video from the beginning at a fixed frame rate (1 FPS by default, adjustable via the API) and built its answer from that sequence of stills[1]. That is a straightforward design, but for a 90-minute lecture or a multi-hour recording, most of the context window ends up filled with frames that have nothing to do with the answer.

Agentic video understanding reverses the order. It pairs the model's reasoning with the native video tools Gemini already has, letting the model work backward from the goal to decide what to watch, at what speed, and through which modality — frames, audio, or transcript. Only the moments and signals it actually needs get pulled in[1].

Under the hood, an agentic loop runs in which the model invokes an internal tool to load just the relevant portion of the video file. Developers could already build this by hand; what has changed is that the assembly work has moved into the model[1].

The groundwork was laid by Agentic Vision, which arrived in Gemini 3 Flash in January. There, the model manipulates images through code execution — zooming, cropping, annotating — to ground its answers in visual evidence, and Google reported a consistent 5–10% quality boost across most vision benchmarks[2]. The same idea has now been extended to video, which adds a time axis.

Cheaper, and still more accurate

Google cites three effects. On standard video analysis benchmarks, analysis costs drop by up to 66%, token consumption by up to 88%, and accuracy rises by up to 7%[1].

The notable part is that cutting cost did not come at the price of accuracy. Irrelevant segments are never loaded, so the input shrinks; segments worth a closer look can be resampled at a higher frame rate, so there is more to reason from.

The gap shows up most clearly on long-form video. From 10-minute how-to guides to 90-minute lectures and multi-hour recordings, developers previously had to choose between paying high token costs and using techniques that drop critical details[1].

Three models support the feature: Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Of these, 3.7 Flash with agentic understanding offers the best overall quality and the best balance of quality against cost, placing it at the accuracy-to-cost pareto frontier among the models tested[1].

Four places where it pays off

The use cases Google highlights are all areas where the old fixed-frame-rate approach struggled[1].

  • Sub-second moment retrieval: pinpoint split-second state changes and tight cut boundaries that 1 FPS misses, making precise automated video editing possible
  • Long-form needle-in-a-haystack search: answer complex queries across multi-hour videos without consuming millions of tokens
  • Anomaly detection: resample interesting time windows at a higher frame rate to inspect rapid motion and subtle visual artifacts
  • Counting actions and objects: accurately track repeated physical movements and distinct objects over time

Machine-reading video has long been designed around a binary choice: show the model everything, or give up. Once the model decides its own viewing scope, that assumption changes.

Turning it on takes one setting

Agentic video understanding is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. It covers both video uploads and YouTube videos, carries no additional feature fee, and uses standard Gemini API token pricing[1].

To enable it, set processing to agentic in the API configuration[1].

from google import genai
client = genai.Client()
interaction = client.interactions.create(
    model="gemini-3.7-flash",
    input=[
        {
            "type": "video",
            "uri": "https://youtu.be/7Z5Vy9JBANs",
            "processing": "agentic"
        },
        {
            "type": "text",
            "text": "What are the 3 most important announcements in this keynote?",
        },
    ],
)
print(interaction.output_text)

Since it is a single line added to existing code, developers already working with long-form video can simply run their current workload and watch what happens to the bill.

Coming to the Gemini app and YouTube answers

This is not limited to the developer API. Google says the feature will roll out to all users in the Gemini app across Flash and Flash-Lite models, and that within the coming months it will also power YouTube's Ask YouTube feature on the video watch page, with the goal of delivering higher-quality answers grounded in the visuals[1].

For a feature like Ask YouTube, used by a broad audience, the cost per question directly determines how much can be offered. If token consumption falls by nearly an order of magnitude, question answering over long videos becomes a realistic option where the economics previously did not work. That is why this announcement is about more than a new API flag.

Summary

Gemini's agentic video understanding shifts video from being read whole at a fixed interval to having the model select the moments and signals it needs. On standard benchmarks, token consumption drops by up to 88%, costs by up to 66%, and accuracy rises by up to 7%. It is supported on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, and can be enabled with a single API setting at no extra charge. The Gemini app and YouTube's Ask YouTube are next in line.

Source [1]: https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-agentic-video-in-gemini/

Source [2]: https://blog.google/innovation-and-ai/technology/developers-tools/agentic-vision-gemini-3-flash/