xAI has released Grok Voice Transcribe 2.0, a speech-to-text API that nearly doubles recognition accuracy compared to its predecessor while keeping pricing unchanged. Independent benchmarks reportedly placed it at the top among major streaming speech recognition models, with notable gains in phone-call audio and short-phrase recognition.
Short-Phrase Errors Cut to Less Than a Third
In tests covering short phrases across 19 languages, the word error rate dropped from 20.6 percent to 6.8 percent. Accuracy improved across categories including customer support calls, everyday conversation, and spoken credentials, with the short-phrase category showing the largest gains. In tests simulating 8kHz telephone audio, the model reportedly outperformed every other model tested on English customer support calls. On an independent benchmark site's ranking of streaming speech recognition models, it placed first among 32 models.
Pricing Stays at a Few Cents an Hour
Pricing remains unchanged from the previous version: 0.10 USD per hour (about 16 yen) for batch processing and 0.20 USD per hour (about 31 yen) for real-time streaming. ※1 USD = 157 JPY (as of September 19, 2026) This pricing already includes features such as speaker diarization, timestamps, and keyword correction.
Speaker Diarization and Multichannel Support
The new model includes speaker diarization to identify who is speaking, word-level timestamps, and the ability to process up to eight audio channels simultaneously. A keyword-biasing feature that improves recognition of specific proper nouns and technical terms supports up to 100 registered terms. It also automatically formats transcribed text for readability, removes filler words characteristic of spoken language, and detects natural pauses to determine where utterances end. The model supports dozens of languages, with automatic language detection and the ability to detect language switching mid-conversation within a single pass. Automatic text formatting is supported in 25 languages. For batch processing, files up to 500MB are supported across 12 audio formats.
API-Only Access, With the Previous Model Retiring Soon
Grok Voice Transcribe 2.0 is available only through the API; there is no open-weight version, and it cannot be self-hosted. Atlassian's Loom video messaging service has already adopted the model, reportedly finding it more accurate than its existing speech recognition solution. The previous version, Grok Voice Transcribe 1.0, is scheduled to be retired within a few weeks, though it will remain accessible in the meantime by specifying its model ID.
Summary
xAI has released Grok Voice Transcribe 2.0, significantly improving recognition accuracy for short phrases and telephone audio while keeping pricing at the same level as the previous version. With added features aimed at practical use cases, including speaker diarization, multichannel support, and keyword biasing, it remains to be seen how widely the model will be adopted for use cases such as call centers and meeting recordings.
