xAI retired its speech-to-text model grok-voice-transcribe-1.0 on October 2, 2026. According to the xAI API release notes, the model “is deprecated and reaches end of life on October 2, 2026,” and all requests to that model name are now routed to grok-voice-transcribe-2.0. The company, which now publishes under the SpaceXAI name, says the newer model is more accurate.

For developers, the practical point is that nothing breaks: code that still sends grok-voice-transcribe-1.0 keeps getting transcripts, but they now come from the 2.0 model. Anyone who pinned 1.0 to keep output stable should rerun their checks, because word timings, formatting and error patterns can differ between two models.

On this page

What changed, and when

Date Change Source
September 17, 2026 grok-voice-transcribe-2.0 added to the Speech-to-Text API; it becomes the default, with 1.0 still selectable xAI release notes
September 18, 2026 xAI publishes its Grok Voice Transcribe 2.0 announcement xAI news
October 2, 2026 grok-voice-transcribe-1.0 reaches end of life; requests to it are routed to 2.0 xAI release notes

The window between the 2.0 release and the 1.0 shutdown was about two weeks.

What 2.0 brings, according to xAI

In its September 18 announcement, xAI says Grok Voice Transcribe 2.0 is “twice as accurate as Grok Voice Transcribe 1.0” across its own real-world evaluations. It is built on the audio foundation model behind Grok Voice. These are the company’s own claims and measurements:

  • Internal tests. xAI measures word error rate on four sets drawn from production traffic: customer-support phone calls, conversations with Grok, spoken credentials such as phone numbers and email addresses, and short voice commands in 19 languages. It says 2.0 improves on 1.0 on all four.
  • Multilingual. xAI calls multilingual accuracy the largest improvement over 1.0. On its short-phrase set, it reports word error rate falling from 20.6% to 6.8%. The model detects the language automatically and follows switches mid-recording.
  • Public leaderboard. xAI says the model ranks first for accuracy among 32 streaming models on the Artificial Analysis leaderboard.

The feature list in the announcement covers batch and streaming transcription, word-level timestamps with confidence scores, speaker diarization, transcription of up to 8 channels, biasing toward up to 100 domain terms per request, written-form formatting of numbers, dates and email addresses, filler-word removal and end-of-turn detection for voice agents. xAI says existing Speech-to-Text API integrations get the accuracy change without code changes.

What developers should check

  • Requests naming grok-voice-transcribe-1.0 keep working but are served by 2.0. Switching the model name to grok-voice-transcribe-2.0 makes the choice explicit in code and logs.
  • Pipelines that parse transcripts, such as call-center analytics or subtitle timing, are worth retesting against the new output.

Retired and current models from xAI and other labs are listed in our AI model release timeline, with the Grok Voice Transcribe 2.0 entry alongside. Other releases and retirements are in AI model news.