
Grok Voice Transcribe 2.0: a 2.7% word error rate at $0.20 an hour, while the API default still ships v1
xAI's Grok Voice Transcribe 2.0 posts the lowest word error rate on Artificial Analysis's streaming leaderboard at an unchanged $0.20 an hour, while requests that omit the model name still route to the version it replaces.
xAI released Grok Voice Transcribe 2.0 on September 18, 2026, its newest speech-to-text model, and describes it as twice as accurate as the version it replaces, at the same price. 1 The model takes recorded files or a live audio stream, and it is built on the audio foundation model behind Grok Voice, the stack xAI also runs in Tesla vehicles. 1
What launched
| Signal | Confirmed detail | Action window |
|---|---|---|
| Model | grok-voice-transcribe-2.0, live for batch files and for real-time streaming. 2 | Nothing to install; it is a hosted endpoint. |
| Price | Unchanged from 1.0: $0.10 an hour of audio in batch and $0.20 an hour streaming, with diarization, timestamps and key terms included. 1 | Re-price a batch backlog. |
| Default slug | Requests that omit model still route to grok-voice-transcribe-1.0; 2.0 becomes the default soon and 1.0 is deprecated in the coming weeks. 3 | Pin the slug you want before it flips. |
| Multilingual | On xAI's short-phrase set across 19 languages, word error rate falls from 20.6% to 6.8%; xAI calls multilingual accuracy its largest gain over 1.0. 1 | Re-test the languages you actually ship. |
| Limits | Files up to 500 MB, 12 audio formats, up to eight channels, and up to 100 bias terms per request. 2 | Check long archives against the ceiling. |
What the numbers say
The accuracy figure you can check independently is on Artificial Analysis's public leaderboard. Grok Voice Transcribe 2.0 records 2.7% on AA-WER Streaming, the lowest bar in the current index, ahead of Muse Voice Transcribe at 3.1% and ElevenLabs Scribe v2 Realtime at 3.6%. The index averages three datasets — AA-AgentTalk, VoxPopuli and Earnings22, about eight hours of audio. 4
The same leaderboard prices Grok Voice Transcribe 2.0 at $3.30 per 1,000 minutes of audio, against $6.50 for ElevenLabs Scribe v2 Realtime and Deepgram Flux, and $1.40 for Inworld STT 1 Realtime. 4

The rest of the evidence is xAI's. Its four internal sets — customer-support telephony, conversations with Grok, spoken credentials and short voice commands — all improve on 1.0; xAI reports 7.2% against 10.6% on the 8 kHz telephony set. Those sets come from xAI's own production traffic, so they measure the audio the model was built for. xAI published no latency figure with the launch. 1
Switching over
Existing integrations take the accuracy gain with no code changes, and the rate card does not move. 1 The trap sits in the default: a pipeline that never pinned the slug keeps transcribing with 1.0 until xAI flips it. 3 Pin
grok-voice-transcribe-2.0 to move, or pin grok-voice-transcribe-1.0 to hold a baseline still while you run both over the same recordings. 2Why it matters
Atlassian put the model into production rather than a pilot: Loom transcribes every recording with it after finding it more accurate than the service it used before, and Atlassian's Sanchan Saxena describes piping a Loom transcript into Cursor to write the code changes. 1
For a call centre or a voice agent, the pair of numbers is the news. The lowest word error rate on the public streaming index arrives at $3.30 per thousand minutes, where ElevenLabs Scribe v2 Realtime and Deepgram Flux each sit at $6.50. Muse Voice Transcribe and Inworld STT 1 Realtime are the cheaper options on the same list, at $3.00 and $1.40. 4
References
- 1
- 2Speech to Text
docs.x.ai
- 3Release notes
docs.x.ai
- 4Speech to Text leaderboard
artificialanalysis.ai
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.