DeepSeek-V4-Flash

Preview

Smaller/cheaper V4 model (~284B total / ~13B active params per authoritative third-party reports). Context length 1M, max output 384K tokens. Input price $0.14/M cache-miss, $0.0028/M cache-hit; output $0.28/M (USD). Supports dual modes (Thinking / Non-Thinking), JSON output, tool calls, chat prefix completion; FIM completion is non-thinking-mode only. Concurrency limit 2500. The legacy aliases deepseek-chat and deepseek-reasoner currently route to this model (non-thinking / thinking respectively). Part of the 'DeepSeek V4 Preview' generation (released 2026-04-24), hence status=preview. Knowledge cutoff NOT officially published by DeepSeek -> left null. Prices re-verified to the cent against api-docs.deepseek.com/quick_start/pricing/ on 2026-08-09. VENDOR-DECLARED FORWARD PRICE EVENT, stated as footnote (2) on that page and still present 2026-08-09: 'We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.' DeepSeek states NO date and NO figure, so no changelog entry and no rate change is recorded here - only the announcement itself. Sourced absence (2026-08-09): the page states no peak/off-peak differential; a 2x off-peak note read on this same page on 2026-08-01 is no longer present, and DeepSeek published no notice of its removal. Responses API supported (footnote (1)). The pricing page is served at the TRAILING-SLASH url (2026-08-11): the extensionless form returns a different document ("Your First API Call") with HTTP 200.

DeepSeek-V4-Flash by DeepSeek costs $0.14 per 1M input tokens and $0.28 per 1M output tokens ($0.175/1M blended), with a 1M (1.000.000-token) context window. It is in preview. Its weights are publicly downloadable, so it can be self-hosted on your own infrastructure.

Last verified: 13 Aug 2026 · sourced from official provider documentation

Provider
DeepSeek
Status
Preview
Input price
$0.14 / 1M tokens
Output price
$0.28 / 1M tokens
Cached input
$0.003 / 1M tokens
Blended price
$0.175 / 1M tokens
Context window
1.000.000 tokens (1M)
Max output
384.000 tokens
Modality
text
Open weights
Yes — self-hostable
Knowledge cutoff
Released
24 Apr 2026
API string
deepseek-v4-flash

Source: DeepSeek official documentation ↗

FAQ

DeepSeek-V4-Flash — questions & answers

How much does DeepSeek-V4-Flash cost?

DeepSeek-V4-Flash is priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens ($0.175/1M blended at a 3:1 input-to-output ratio), with cached input at $0.003 per 1M tokens.

What is the context window of DeepSeek-V4-Flash?

DeepSeek-V4-Flash has a 1.000.000-token context window (1M), with up to 384.000 output tokens per request.

Is DeepSeek-V4-Flash deprecated?

No — DeepSeek-V4-Flash is in preview and not currently scheduled for deprecation or retirement in our tracker.

Does DeepSeek-V4-Flash support image or vision input?

Our catalog lists DeepSeek-V4-Flash's modalities as text; image/vision input is not among them.

Can I self-host DeepSeek-V4-Flash?

Yes — DeepSeek-V4-Flash is an open-weight model: its weights are publicly downloadable, so you can run it on your own infrastructure (subject to its license) instead of only calling DeepSeek's API.

What is the API model string for DeepSeek-V4-Flash?

The API model identifier for DeepSeek-V4-Flash is "deepseek-v4-flash" when calling DeepSeek's API.