Dynamic Beating AI News: Google has launched Agentic Video Understanding on the Gemini API. It allows the model to decide which parts of a video to watch and how to analyze them. The default standard mode samples frames at a fixed rate of 1 FPS and processes them all at once within the context. The new mode autonomously identifies segments along the timeline, selects imagery, audio, or transcript text based on the query, and can reevaluate rapidly moving actions by increasing the frame rate.
Google's internal testing reports that with Gemini 3.7 Flash enabled, Token consumption decreased by up to 88%, analysis costs decreased by up to 66%, and relative accuracy increased by approximately 7%. The system can pinpoint moments lasting less than a second and identify specific details within hours-long videos.
Technically, this feature integrates the "locate first, then dive deeper" workflow that developers previously had to set up themselves into the Gemini platform. Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite all now support this functionality. It is applicable to uploaded videos and YouTube content at no additional cost. Future integrations will include the Gemini App and YouTube's Ask YouTube feature.

