Gemini 3.5 Transcribe
GASpeech-to-text. Audio in, text plus word annotations out, with automatic detection of 85+ languages, mid-session code-switching, speaker diarization (up to 8 speakers), word-level timestamps, smart dictation and formatting, and custom vocabulary biasing of up to 1,000 terms. Status ga because Google's models index badges the Gemini 3.5 Transcribe card 'New Stable' — one card covering both endpoints, gemini-3.5-transcribe and gemini-3.5-transcribe-live — and the id carries no -preview suffix; read 2026-09-01. Paid-tier rates per 1M tokens: input $2.00 on audio, output $12.00 on text. Google also publishes duration-equivalent figures for the same rates (roughly $0.003 per minute of audio in and $0.002 per minute of text out, on an assumption of 25 audio tokens per second and 175 text tokens per minute, blending to about $0.005 per minute), which this schema has no field for. Free tier is free of charge. No context window or max output token count is published for this endpoint, so both fields are null rather than inherited from a sibling; the stated limit is duration — up to 1 hour of audio per request, dropping to 30 minutes when diarization or word-level timestamps are enabled. Caching, code execution, file search, function calling, image generation and thinking are all listed Not supported, as are the batch, flex and priority consumption options. Released August 2026 per Google's deprecations table — month only, which is all Google states, so the field carries 2026-08 rather than a guessed day — and that table says 'No shutdown date announced' with the row not gray. Added 2026-09-01: missing from the catalog since launch, one of four Google models invisible to the check-google-prices.mjs completeness column because all of its price cells are multi-rate (LEARNINGS #105).
Gemini 3.5 Transcribe by Google costs $2 per 1M input tokens and $12 per 1M output tokens ($4.50/1M blended). It is generally available (GA).
Last verified: 26 Sep 2026 · sourced from official provider documentation
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · no spam · unsubscribe anytime.
Source: Google official documentation ↗
Gemini 3.5 Transcribe — questions & answers
How much does Gemini 3.5 Transcribe cost?
Gemini 3.5 Transcribe is priced at $2 per 1M input tokens and $12 per 1M output tokens ($4.50/1M blended at a 3:1 input-to-output ratio).
Is Gemini 3.5 Transcribe deprecated?
No — Gemini 3.5 Transcribe is generally available (GA) and not currently scheduled for deprecation or retirement in our tracker.
Does Gemini 3.5 Transcribe support image or vision input?
Our catalog lists Gemini 3.5 Transcribe's modalities as audio, text; image/vision input is not among them.
Can I self-host Gemini 3.5 Transcribe?
No — Gemini 3.5 Transcribe is proprietary. Its weights are not publicly released, so it is available only through Google's API (or hosted partners), not for self-hosting.
What is the API model string for Gemini 3.5 Transcribe?
The API model identifier for Gemini 3.5 Transcribe is "gemini-3.5-transcribe" when calling Google's API.
Track Gemini 3.5 Transcribe price & status changes
New models, price cuts, and deprecations — a short email when something actually changes. No spam, unsubscribe anytime.
The last few changes we caught — this is what lands in your inbox:
- DeprecatingGoogle re-pointed the Gemini 2.5 TTS migration path past the 3.1 preview and onto the new GA 3.8 TTS models
- New modelGemini 3.8 Flash TTS is GA - a flagship text-to-speech model at under half the audio rate of the preview it replaces, on a two-column price schedule
- New modelGemini 3.8 Flash-Lite TTS is GA - Google's cost-efficient TTS, which Google says was built to replace Gemini 3.1 Flash TTS
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · powered by the same data on this page.