Gemini 2.0 Flash-Lite

Retired

RETIRED/shut down 2026-06-01 per Google's deprecations table (released 2025-02-25, replacement gemini-3.1-flash-lite; gemini-2.0-flash-lite-001 shares both dates). Re-verified first-party 2026-07-31: Google keeps shut-down models in BOTH the live pricing page and a dedicated model page, so every field here is now read, not inferred. Pricing page, paid tier: $0.075 in / $0.30 out per 1M tokens (batch tier $0.0375 / $0.15 — half, but no schema field for it yet), "Context caching price: Not available" on both tiers. Model page confirms inputs audio/images/video/text, 1,048,576 input tokens, 8,192 output tokens, knowledge cutoff August 2024. Note the two surfaces differ on caching: the model page lists Caching as Supported while the pricing page publishes no cached-input rate, so price_cached_input_per_mtok stays null — the capability existed, the price was never published (an earlier note here wrongly said caching was not offered).

Gemini 2.0 Flash-Lite by Google costs $0.075 per 1M input tokens and $0.3 per 1M output tokens ($0.131/1M blended), with a 1.0M (1.048.576-token) context window. It has been retired (shut down 1 Jun 2026); the recommended migration is Gemini 3.1 Flash-Lite.

Last verified: 13 Aug 2026 · sourced from official provider documentation

Lifecycle warning

This model has been retired. Scheduled shutdown: 1 Jun 2026. Recommended migration: Gemini 3.1 Flash-Lite.

Read the Gemini 2.0 Flash-Lite migration guide → replacement & cheapest alternatives

Provider
Google
Status
Retired
Input price
$0.075 / 1M tokens
Output price
$0.3 / 1M tokens
Cached input
Blended price
$0.131 / 1M tokens
Context window
1.048.576 tokens (1.0M)
Max output
8192 tokens
Modality
text, image, video, audio
Open weights
No — API only
Knowledge cutoff
2024-08
Released
25 Feb 2025
API string
gemini-2.0-flash-lite

Source: Google official documentation ↗

FAQ

Gemini 2.0 Flash-Lite — questions & answers

How much does Gemini 2.0 Flash-Lite cost?

Gemini 2.0 Flash-Lite is priced at $0.075 per 1M input tokens and $0.3 per 1M output tokens ($0.131/1M blended at a 3:1 input-to-output ratio).

What is the context window of Gemini 2.0 Flash-Lite?

Gemini 2.0 Flash-Lite has a 1.048.576-token context window (1.0M), with up to 8192 output tokens per request.

Is Gemini 2.0 Flash-Lite still available?

Gemini 2.0 Flash-Lite has been retired and is no longer available. Its scheduled shutdown date is 1 Jun 2026. The recommended migration is Gemini 3.1 Flash-Lite.

Does Gemini 2.0 Flash-Lite support image or vision input?

Yes — Gemini 2.0 Flash-Lite accepts image input. Listed modalities: text, image, video, audio.

Can I self-host Gemini 2.0 Flash-Lite?

No — Gemini 2.0 Flash-Lite is proprietary. Its weights are not publicly released, so it is available only through Google's API (or hosted partners), not for self-hosting.

What is the API model string for Gemini 2.0 Flash-Lite?

The API model identifier for Gemini 2.0 Flash-Lite is "gemini-2.0-flash-lite" when calling Google's API.

What is Gemini 2.0 Flash-Lite's knowledge cutoff?

Gemini 2.0 Flash-Lite's training knowledge cutoff is Aug 2024.