●VIDEO — Agentic video understanding reached 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite on September 1. The model navigates the timeline itself rather than sampling frames at a fixed rate●TOKENS — Because it pulls transcripts, frames, or audio only when it needs them, Google measures up to 88% fewer tokens on long-form content●SCOPE — It works across both the Interactions and GenerateContent APIs. If you have costed out long-video work before, the assumptions have moved●MUSIC — Lyria 3.5 entered public preview on September 3, generating full-length songs at 44.1 kHz stereo●CONTROL — Lyria 3.5 accepts text and image inputs, with better musical coherence, more natural vocals, and finer control over duration and structure●ROBOTICS — gemini-robotics-er-2-streaming-preview is tuned for real-time streaming over the Live API, with function calling that blocks on physical robot actions●VIDEO — Agentic video understanding reached 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite on September 1. The model navigates the timeline itself rather than sampling frames at a fixed rate●TOKENS — Because it pulls transcripts, frames, or audio only when it needs them, Google measures up to 88% fewer tokens on long-form content●SCOPE — It works across both the Interactions and GenerateContent APIs. If you have costed out long-video work before, the assumptions have moved●MUSIC — Lyria 3.5 entered public preview on September 3, generating full-length songs at 44.1 kHz stereo●CONTROL — Lyria 3.5 accepts text and image inputs, with better musical coherence, more natural vocals, and finer control over duration and structure●ROBOTICS — gemini-robotics-er-2-streaming-preview is tuned for real-time streaming over the Live API, with function calling that blocks on physical robot actions