My conversation with @koraykv, the lead for Google DeepMind, about the path to AGI, progress with 3.7 Flash, why we are laser focused on the frontier, and more!
Excited to bring Gemini Omni Flash 1.1 to the world, this is an update to our anything in, anything out world model. It now supports:
- 360p drafts, 4K up samplers
- up to 10 seconds of video context when extending
- video extensions in 10 second increments
- video references
Introducing Gemini 3.5 Transcribe, our new speech to text model with smart transcription, function calling, more precise transcription (lower WER), custom vocabulary support, multi-speaker identification, and support for over 85 languages!
Also with realtime streaming support!