Gemini 3.5 Transcribe succeeds Chirp 3 with dual live and file speech APIs

Gemini 3.5 Transcribe succeeds Chirp 3 with dual live and file speech APIs

Google DeepMind's Gemini 3.5 Transcribe is in public preview with separate live and file endpoints, 85+ languages, and company-reported gains over Chirp 3.

Google DeepMind released Gemini 3.5 Transcribe on August 26, 2026, as the successor to Chirp 3 for speech-to-text. It ships as two public-preview endpoints: gemini-3.5-transcribe for pre-recorded audio on the Interactions API, and gemini-3.5-transcribe-live for real-time bidirectional streaming on the Live API. Both are in Google AI Studio and the Gemini API; enterprises can try Live on Gemini Enterprise Agent Platform. 12

What launched

SignalConfirmed detailWhy it matters
Model IDsFile path: gemini-3.5-transcribe. Live path: gemini-3.5-transcribe-live. Status is public preview. 12Two separate contracts; match the endpoint to live voice or batch files.
LanguagesAuto-detects 85+ locales, including mid-session code-switching. 23Multilingual meetings need no fixed language setting.
Smart vs verbatimsmart removes fillers, resolves self-corrections, and formats lists and numbers; verbatim keeps the raw transcript and is required for diarization or word timestamps. 3Dictation wants smart mode; compliance archives want verbatim.
File and live limitsUp to 1 hour of audio per unary request; 30 minutes when diarization or word-level timestamps are on. Live sessions cap at 10 minutes. Diarization supports up to 8 speakers, with 3+ speaker attribution marked experimental. Word timestamps are file-only and may lower accuracy. 234Long multi-speaker meetings need the file API and a split plan.
VocabularyCustom vocabulary accepts up to 1,000 terms; docs say best results are usually with up to 100. 3Brand names and SKUs can be biased without fine-tuning.
PricePaid tier for gemini-3.5-transcribe: $2.00 / $12.00 per 1M tokens (~$0.005/min blended). Live: $3.50 / $21.00 (~$0.009/min). Free tier available. 5Live costs more; batch analytics stay cheaper per minute.
The model card places Gemini 3.5 Transcribe and Transcribe Live on Gemini 3 Pro, with up to 96K input tokens and 32K output tokens. Known limits include hallucinations plus occasional slowness or timeouts; the knowledge cutoff is January 2025. 6

Company-reported accuracy

FLEURS streaming word error rates for Gemini 3.5 Transcribe Live versus Chirp 3 and peer models
Google's FLEURS top-locales streaming chart lists Gemini 3.5 Transcribe Live at 5.50% WER against Chirp 3 at 7.32%, ElevenLabs Scribe v2 Realtime at 9.70%, OpenAI GPT Live Transcribe at 8.97%, and Deepgram Nova-3 at 15.77% (lower is better). 17
On the matching non-streaming FLEURS slice, Google reports 5.04% WER for Gemini 3.5 Transcribe versus 5.66% for Chirp 3. Separately, Google cites Artificial Analysis averages of 2.6% WER for non-streaming and 4.0% for streaming, plus a 70% faster time-to-final-transcription versus Chirp 3. Those AA figures are company-cited third-party scores; the FLEURS bars above are Google's own comparison chart. 17

Why this matters

Chirp 3's successor is now a first-class Gemini API surface with separate live and file paths, smart cleanup for dictation, and hooks already in Gboard Rambler, Antigravity, and the Gemini macOS app. For voice agents, captions, or post-call analytics, run a short preview trial: measure alphanumeric accuracy and latency on gemini-3.5-transcribe-live for interactive turns, then put the same noisy multi-speaker samples through gemini-3.5-transcribe in smart and verbatim modes. Preview status, the 10-minute live session cap, and experimental multi-speaker attribution are the main gates to plan around. 12

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content

More from this channel