Grok Voice Think Fast 2.0

GA

Realtime speech-to-speech model for voice agents, driven over a WebSocket at wss://api.x.ai/v1/realtime with low-latency turn-taking and server-side tool use. Billed by connected audio time and by session concurrency rather than by token, so the token price fields are null by design. Released 2026-07-29, dated by xAI's own release notes ('grok-voice-think-fast-2.0 is now available with Speech to Speech'), and the alias grok-voice-latest was stated in the same entry to route here from 2026-08-05 - so an application pinned to that alias moved to this model on that date without changing a string. THE PUBLISHED RATE IS $0.08 per minute of audio, which xAI itself also prints as $4.80/hr, and the models-index hydration payload agrees: 'realtimeAudioSecondPrice': '13333333' in the 1e-10 USD units this catalog calibrated against rows it already held, which is $0.0013333 per second. The payload additionally carries 'pstnMinutePrice': '100000000' - $0.01 per minute, the same integer this catalog already publishes as $0.01 per image elsewhere in xAI's payload, so the unit is confirmed rather than assumed - for sessions bridged to the public telephone network. TWO FURTHER PAYLOAD FIELDS ARE DELIBERATELY NOT CARRIED, and this is a closed path rather than an oversight: 'realtimeAudioTokenPrice': '266667' and 'realtimeTextInputPrice': '40000000' cannot both be read in the per-token unit that works everywhere else in this payload - the second would come out four orders of magnitude above any published rate - and xAI's pricing page prints the second one as a bare figure of $0.004 for realtime text with no denominator at all, so neither surface establishes what either field is per. A figure whose unit is unknown is not a price, and this catalog will not publish one. Session concurrency is the rate-limit axis: 10 concurrent sessions at the base tier, 20 / 50 / 100 / 200 / 400 at tiers 1 to 5. A first-generation grok-voice-think-fast-1.0 was released 2026-04-23 and is absent from the current models payload; this catalog does NOT infer a retirement from that absence and carries no row for it. Added to the catalog 2026-09-19 as a completeness fix; xAI published nothing on this date.

Grok Voice Think Fast 2.0 by xAI is tracked live in our LLM catalog. It is generally available (GA).

Last verified: 26 Sep 2026 · sourced from official provider documentation

Provider
xAI
Status
GA
Input price
—
Output price
—
Cached input
—
Blended price
—
Context window
—
Max output
—
Modality
text-in, audio, text-out, audio-out
Open weights
No — API only
Knowledge cutoff
—
Released
29 Jul 2026
API string
grok-voice-think-fast-2.0

Source: xAI official documentation ↗

FAQ

Grok Voice Think Fast 2.0 — questions & answers

Is Grok Voice Think Fast 2.0 deprecated?

No — Grok Voice Think Fast 2.0 is generally available (GA) and not currently scheduled for deprecation or retirement in our tracker.

Does Grok Voice Think Fast 2.0 support image or vision input?

Our catalog lists Grok Voice Think Fast 2.0's modalities as text-in, audio, text-out, audio-out; image/vision input is not among them.

Can I self-host Grok Voice Think Fast 2.0?

No — Grok Voice Think Fast 2.0 is proprietary. Its weights are not publicly released, so it is available only through xAI's API (or hosted partners), not for self-hosting.

What is the API model string for Grok Voice Think Fast 2.0?

The API model identifier for Grok Voice Think Fast 2.0 is "grok-voice-think-fast-2.0" when calling xAI's API.