header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Google Grants Gemini Video Agent: Costliest Falls 66%

Dynamic Beating AI News: Google has launched Agentic Video Understanding on the Gemini API. It allows the model to decide which parts of a video to watch and how to analyze them. The default standard mode samples frames at a fixed rate of 1 FPS and processes them all at once within the context. The new mode autonomously identifies segments along the timeline, selects imagery, audio, or transcript text based on the query, and can reevaluate rapidly moving actions by increasing the frame rate.


Google's internal testing reports that with Gemini 3.7 Flash enabled, Token consumption decreased by up to 88%, analysis costs decreased by up to 66%, and relative accuracy increased by approximately 7%. The system can pinpoint moments lasting less than a second and identify specific details within hours-long videos.


Technically, this feature integrates the "locate first, then dive deeper" workflow that developers previously had to set up themselves into the Gemini platform. Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite all now support this functionality. It is applicable to uploaded videos and YouTube content at no additional cost. Future integrations will include the Gemini App and YouTube's Ask YouTube feature.

Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish