Grok Voice Transcribe 2.0
GASpeech-to-text model, billed by duration of audio rather than by token, so the token price fields are null by design. TWO RATES, NOT ONE, AND THE MODE DECIDES WHICH: xAI's pricing page lists Speech to Text at $0.10 per hour (REST) and $0.20 per hour (Streaming) - streaming is exactly double - and the models-index hydration payload carries both as integers in the 1e-10 USD-per-second units this catalog calibrated against rows it already held: 'perAudioSecond': '277778' and 'perAudioSecondStreaming': '555556'. Released 2026-09-17, dated by xAI's own release notes: 'grok-voice-transcribe-2.0 is now available. Use grok-voice-transcribe-1.0 or grok-voice-transcribe-2.0'. THE DEFAULT MODEL IS CONTRADICTED ACROSS TWO FIRST-PARTY SURFACES and this catalog does not pick a winner: that release-notes entry says 'the default is grok-voice-transcribe-1.0', while the Speech to Text reference prints grok-voice-transcribe-2.0 in the model parameter's default cell. Set model explicitly and the question does not arise. Capabilities, from the reference: 25 languages, batch and streaming modes, 12 audio formats, word-level timestamps, multichannel transcription (2-8 channels), speaker diarization, Smart Turn end-of-turn detection, and a vad_threshold parameter that tunes the voice-activity gate (0 disables it). Maximum 500 MB per file, or a URL the server downloads. Rate limits at the base tier are 10 rps / 600 rpm / 100 concurrent sessions. Added to the catalog 2026-09-19, two days after release.
Grok Voice Transcribe 2.0 by xAI is tracked live in our LLM catalog. It is generally available (GA).
Last verified: 26 Sep 2026 · sourced from official provider documentation
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · no spam · unsubscribe anytime.
Source: xAI official documentation ↗
Grok Voice Transcribe 2.0 — questions & answers
Is Grok Voice Transcribe 2.0 deprecated?
No — Grok Voice Transcribe 2.0 is generally available (GA) and not currently scheduled for deprecation or retirement in our tracker.
Does Grok Voice Transcribe 2.0 support image or vision input?
Our catalog lists Grok Voice Transcribe 2.0's modalities as audio, text-out; image/vision input is not among them.
Can I self-host Grok Voice Transcribe 2.0?
No — Grok Voice Transcribe 2.0 is proprietary. Its weights are not publicly released, so it is available only through xAI's API (or hosted partners), not for self-hosting.
What is the API model string for Grok Voice Transcribe 2.0?
The API model identifier for Grok Voice Transcribe 2.0 is "grok-voice-transcribe-2.0" when calling xAI's API.
Track Grok Voice Transcribe 2.0 price & status changes
New models, price cuts, and deprecations — a short email when something actually changes. No spam, unsubscribe anytime.
The last few changes we caught — this is what lands in your inbox:
- DeprecatingGoogle re-pointed the Gemini 2.5 TTS migration path past the 3.1 preview and onto the new GA 3.8 TTS models
- New modelGemini 3.8 Flash TTS is GA - a flagship text-to-speech model at under half the audio rate of the preview it replaces, on a two-column price schedule
- New modelGemini 3.8 Flash-Lite TTS is GA - Google's cost-efficient TTS, which Google says was built to replace Gemini 3.1 Flash TTS
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · powered by the same data on this page.