DeepSeek-V4-Flash

Retired

RETIRED 2026-09-10, superseded by DeepSeek-V4.1-Flash (deepseek-flash). Signal: an explicit past-tense statement by the provider on a dated announcement — DeepSeek's Change Log entry for 2026-09-10 reads 'The previous-generation models V4 Flash and V4 Flash Vision Exp have been retired; for compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily routed to V4.1 Flash', and the pricing page footnote says the same: 'The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but the corresponding models have been retired, their requests are served by the DeepSeek-V4.1-Flash model and billed at the Flash price.' That is a statement, not an inference from the column vanishing off the pricing page, though it vanished the same day. DeepSeek published no separate shutdown date and no deprecation window: one date, so this row carries it in retires_on only and leaves deprecated_on null. THE NAME STILL WORKS — calls do not fail, they are served by V4.1-Flash and billed at the cheaper Flash price, and DeepSeek calls that routing 'temporary', so consumers should migrate to deepseek-flash rather than rely on it. Until 2026-09-10 this row was carried as preview, on the grounds that it was part of the 'DeepSeek V4 Preview' generation released 2026-04-24, and that was right for the period it covered. PRICES ABOVE ARE THE RATES IN FORCE AT RETIREMENT and deliberately no longer track the endpoint that now serves the traffic: PEAK $0.014/M cache-hit, $0.44/M cache-miss, $1.32/M out, with off-peak exactly half. V4.1-Flash is cheaper on every axis ($0.006 / $0.30 / $1.20 peak), so a caller left on the legacy name is billed at the new, lower Flash price, not at these. Historical spec, unchanged: smaller/cheaper V4 model (~284B total / ~13B active params per authoritative third-party reports), 1M context, 384K max output, dual Thinking / Non-Thinking modes, JSON output, tool calls, chat prefix completion, FIM completion in non-thinking mode only, Responses API and Anthropic-format API supported, concurrency limit 2500, open weights. Knowledge cutoff never published by DeepSeek, so it stays null. The legacy aliases deepseek-chat and deepseek-reasoner pointed here and were themselves retired 2026-07-24, which makes that pair a two-hop chain to deepseek-v4-1-flash; each row records the target its provider named at the time rather than being silently re-pointed. Prior lifecycle of this row, kept because it explains the price fields: DeepSeek moved API billing to peak/off-peak at 16:00 UTC on 2026-08-16, having announced it undated on 2026-08-13 and dated on 2026-08-14; the fields were re-pointed at the PEAK column on the cutover day, about 8 hours early, because this catalog refreshes once daily at ~08:05 UTC and the alternative was publishing superseded rates for ~16 hours. That window closed on 2026-08-17.

DeepSeek-V4-Flash by DeepSeek costs $0.44 per 1M input tokens and $1.32 per 1M output tokens ($0.66/1M blended), with a 1M (1.000.000-token) context window. It has been retired (shut down 10 Sep 2026); the recommended migration is DeepSeek-V4.1-Flash. Its weights are publicly downloadable, so it can be self-hosted on your own infrastructure.

Last verified: 26 Sep 2026 · sourced from official provider documentation

Retired on 10 Sep 2026

Lifecycle warning

This model has been retired. Recommended migration: DeepSeek-V4.1-Flash.

Read the DeepSeek-V4-Flash migration guide → replacement & cheapest alternatives

Provider
DeepSeek
Status
Retired
Input price
$0.44 / 1M tokens
Output price
$1.32 / 1M tokens
Cached input
$0.014 / 1M tokens
Blended price
$0.66 / 1M tokens
Context window
1.000.000 tokens (1M)
Max output
384.000 tokens
Modality
text
Open weights
Yes — self-hostable
Knowledge cutoff
—
Released
24 Apr 2026
API string
deepseek-v4-flash

Source: DeepSeek official documentation ↗

FAQ

DeepSeek-V4-Flash — questions & answers

How much does DeepSeek-V4-Flash cost?

DeepSeek-V4-Flash is priced at $0.44 per 1M input tokens and $1.32 per 1M output tokens ($0.66/1M blended at a 3:1 input-to-output ratio), with cached input at $0.014 per 1M tokens.

What is the context window of DeepSeek-V4-Flash?

DeepSeek-V4-Flash has a 1.000.000-token context window (1M), with up to 384.000 output tokens per request.

Is DeepSeek-V4-Flash still available?

DeepSeek-V4-Flash has been retired and is no longer available. Its scheduled shutdown date is 10 Sep 2026. The recommended migration is DeepSeek-V4.1-Flash.

Does DeepSeek-V4-Flash support image or vision input?

Our catalog lists DeepSeek-V4-Flash's modalities as text; image/vision input is not among them.

Can I self-host DeepSeek-V4-Flash?

Yes — DeepSeek-V4-Flash is an open-weight model: its weights are publicly downloadable, so you can run it on your own infrastructure (subject to its license) instead of only calling DeepSeek's API.

What is the API model string for DeepSeek-V4-Flash?

The API model identifier for DeepSeek-V4-Flash is "deepseek-v4-flash" when calling DeepSeek's API.