Models

Google Ships Gemini 3.5 Transcribe and Omni 1.1 Flash

Gemini 3.5 Transcribe turns audio into formatted text in 85+ languages at a 2.6 percent error rate; Omni 1.1 Flash extends video to 40 seconds and adds a draft mode at a third of the cost.

Faruk TalmaçAugust 29, 20263 min read4 views
Google Ships Gemini 3.5 Transcribe and Omni 1.1 Flash

Google shipped two specialist models in two days, and the more useful one for most businesses is the less glamorous one. Gemini 3.5 Transcribe, released in public preview on August 26, turns raw audio into punctuated, formatted text with the filler words removed, across more than 85 automatically detected languages. Gemini Omni 1.1 Flash, released to developers on August 27, is a video generation model whose headline feature is control rather than spectacle.

Transcribe: the numbers

Google reports a word error rate of 2.6 percent on recorded audio and 4.0 percent in streaming mode. On the multilingual FLEURS benchmark the figures are 5.04 and 5.50 percent respectively. Latency is down 70 percent compared with Chirp 3, the model it replaces. Language detection is automatic across the 85-plus supported languages, with handling for regional accents and dialects.

The feature list is aimed squarely at meetings and calls: speaker identification for up to three voices, word-level timestamps, custom vocabulary for product names and jargon, and cleanup of disfluencies and self-corrections so the transcript reads the way the speaker meant it. It is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, and it already powers Gboard's "Rambler" feature on Android and the Gemini app on macOS, with Chrome next. Google did not state a price in the announcement; the API pricing table bills by token, and third-party estimates put it well under a cent per minute, a figure worth checking against your own usage.

Omni 1.1 Flash: longer scenes, cheaper drafts

The video model can now extend a generated clip in 10-second increments up to 40 seconds in total, and it looks at the previous 10 seconds of footage while doing so. The earlier version referenced only the final second, which is why extensions tended to drift. You can also fix the first and last frame of a scene, reference up to three seconds of existing video, and generate in a 360p draft mode that is up to 60 percent faster and a third of the cost of standard 720p, then upscale the keeper to 1080p or 4K. It is available in AI Studio, the Enterprise Agent Platform API, Google Flow and the Gemini app; Adobe, Figma Weave, Runway and GMI Cloud have already integrated it.

Why Transcribe is the one to test first

Video generation is still a creative-team tool. Transcription touches every company that records anything: sales calls, board meetings, support lines, field reports, clinical notes. Most of that audio is currently either typed up by a person or run through an older engine that stumbles on accents and mixed languages. A model that detects the language, separates speakers and drops the "ums" removes the cleanup step that made transcripts too expensive to bother with. The sensible move is a 30-minute pilot on your own recordings, paying particular attention to how speaker separation behaves in your language and setting, before deciding whether meeting notes and call summaries can be automated end to end.

Sources: Google (Gemini 3.5 Transcribe), Google (Gemini Omni 1.1 Flash)

Share This Article

Faruk Talmaç

Written by

Faruk Talmaç

Co-Founder & Editor

Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.

More news

Want to put this technology to work in your business?

Let's talk