
Introducing agentic video understanding with Gemini
Google launched agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. This feature replaces static frame processing with dynamic tools to improve analysis accuracy and efficiency.
Why it matters
Developers can analyze long-form videos more accurately while significantly reducing token costs. This enables high-precision automated video editing and efficient searching of multi-hour recordings.
The details
The feature supports sub-second moment retrieval, anomaly detection, and precise object counting. It works with both video uploads and YouTube videos. There is no additional feature fee beyond standard API token pricing. Gemini dynamically chooses between frames, audio, and transcripts based on the goal.
What's next
The feature will roll out to the Gemini app and power YouTube's 'Ask YouTube' feature in coming months.
Show entities and relationshipsHide entities and relationships
In this article
Companies
Key connections
agentic video understanding is related to Anomaly Detection
unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection
Gemini 3.7 Flash uses agentic video understanding
Activating agentic video understanding drops token consumption by up to 88% and boosts accuracy by up to 7% with Gemini 3.7 Flash.
agentic video understanding is related to agentic vision
Similar to agentic vision, which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.