Gemini 3.5 Flash-Lite

GA

GA/stable — 'Stable: gemini-3.5-flash-lite' on its model page, described by Google as 'our most cost-efficient GA model, optimized for high-volume agentic tasks, translation, and simple data processing'. Added 2026-08-02 because Google's deprecations table names it the recommended replacement for gemini-3.1-flash-lite (shutdown 2027-05-07). Prices are the paid-tier STANDARD rates: $0.30 in / $2.50 out per 1M, context caching $0.03/1M plus a $1.00 per 1M-tokens-per-hour storage fee this schema has no field for. Unlike Gemini 3.1 Flash-Lite and 2.5 Flash-Lite, there is NO higher audio-input tier — Google publishes one input rate covering text/image/video/audio. Batch tier $0.15 / $1.25. Grounding with Google Search or Maps: 5,000 requests/month free shared across all Gemini 3.x models, then $14 per 1,000. Spec table gives 1,048,576 input / 65,536 output tokens, inputs text/image/video/audio/PDF, output text only; caching, code execution, function calling, file search and Google Maps grounding supported, computer use in preview, no Live API and no image or audio generation. Release date July 21, 2026 read from Google's deprecations table and corroborated by the API changelog entry of the same date ('Gemini 3.6 Flash and Gemini 3.5 Flash-Lite generally available'); no shutdown date announced. Google publishes no knowledge cutoff for it on the model page, so that field is null rather than inherited from a sibling — the page's 'Latest update July 2026' is a docs-edit date, not a cutoff.

Gemini 3.5 Flash-Lite by Google costs $0.3 per 1M input tokens and $2.50 per 1M output tokens ($0.85/1M blended), with a 1.0M (1.048.576-token) context window. It is generally available (GA).

Last verified: 13 Aug 2026 · sourced from official provider documentation

Provider
Google
Status
GA
Input price
$0.3 / 1M tokens
Output price
$2.50 / 1M tokens
Cached input
$0.03 / 1M tokens
Blended price
$0.85 / 1M tokens
Context window
1.048.576 tokens (1.0M)
Max output
65.536 tokens
Modality
text, image, video, audio, pdf
Open weights
No — API only
Knowledge cutoff
Released
21 Jul 2026
API string
gemini-3.5-flash-lite

Source: Google official documentation ↗

FAQ

Gemini 3.5 Flash-Lite — questions & answers

How much does Gemini 3.5 Flash-Lite cost?

Gemini 3.5 Flash-Lite is priced at $0.3 per 1M input tokens and $2.50 per 1M output tokens ($0.85/1M blended at a 3:1 input-to-output ratio), with cached input at $0.03 per 1M tokens.

What is the context window of Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite has a 1.048.576-token context window (1.0M), with up to 65.536 output tokens per request.

Is Gemini 3.5 Flash-Lite deprecated?

No — Gemini 3.5 Flash-Lite is generally available (GA) and not currently scheduled for deprecation or retirement in our tracker.

Does Gemini 3.5 Flash-Lite support image or vision input?

Yes — Gemini 3.5 Flash-Lite accepts image input. Listed modalities: text, image, video, audio, pdf.

Can I self-host Gemini 3.5 Flash-Lite?

No — Gemini 3.5 Flash-Lite is proprietary. Its weights are not publicly released, so it is available only through Google's API (or hosted partners), not for self-hosting.

What is the API model string for Gemini 3.5 Flash-Lite?

The API model identifier for Gemini 3.5 Flash-Lite is "gemini-3.5-flash-lite" when calling Google's API.