DeepSeek-V4.1-Flash
GADeepSeek's first model on its new architecture family, released 2026-09-10 and dated first-hand by DeepSeek's own Change Log entry for that day: 'Today, we officially release the DeepSeek-V4.1-Flash model. It is the smallest model in our new architecture family, with native multimodal visual understanding.' Call it as deepseek-flash — that is the model name DeepSeek instructs you to use, and it is why api_string is not the row id. PRICES ARE A PEAK / OFF-PEAK PAIR, so read the column, not the number: the fields above carry the PEAK rates per 1M tokens ($0.30 cache-miss input, $1.20 output, $0.006 cache-hit input), and off-peak is exactly half on every axis ($0.15 / $0.60 / $0.003). DeepSeek states, as revised on or just before 2026-09-19, 'Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday, excluding Chinese public holidays. All other hours are off-peak, including weekends and Chinese public holidays in full.' - so off-peak is what you pay for most of the week, and on a Chinese public holiday you pay it around the clock. Until that revision the same footnote read 'Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak)', with no holiday carve-out; this catalog's prose baseline carried the old wording on 2026-09-18 and the new wording on 2026-09-19, and DeepSeek dated the edit nowhere. No rate moved. The fields track peak because DeepSeek presents peak as the rate and off-peak as half of it, which is the same convention this catalog has applied to the V4 rows since the 2026-08-16 cutover. That makes this a genuine price cut for the Flash line, and DeepSeek says so in the same Change Log entry ('With the release of DeepSeek-V4.1-Flash, API prices have been reduced accordingly'): the model it replaces, deepseek-v4-flash, was $0.44 / $1.32 / $0.014 at peak, so input falls ~32%, output ~9% and cache-hit input ~57%. Native vision: the feature table marks Vision supported for this model and 'Not supported' for deepseek-v4-pro, and the HuggingFace repo is tagged image-text-to-text, hence modality text + image. Thinking and non-thinking modes (thinking is the default), JSON output, tool calls, Responses API, Anthropic-format API and chat prefix completion all supported; FIM completion is non-thinking-mode only. Context length 1M, max output 384K, concurrency limit 2500. Open weights published the same day at huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash under the MIT license (fp8/8-bit safetensors). Status is ga because DeepSeek announces it as an official release with no preview or experimental label — unlike the V4 generation, which it shipped as 'DeepSeek V4 Preview' and which this catalog therefore carried as preview. Knowledge cutoff is NOT published by DeepSeek, so the field is null rather than guessed. It is the stated destination for two retired endpoints: deepseek-v4-flash and deepseek-v4-flash-vision-exp both route here, their names still accepted but the models behind them retired. Until 2026-09-12 this note named a third, deepseek-v4-pro, as joining them at 12:00 Beijing Time on 2026-09-14; DeepSeek has since withdrawn that retirement - 'we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged' - so V4 Pro stays live on its own rates and does not route here. The pricing page is served at the TRAILING-SLASH url; the extensionless form returns a different document ('Your First API Call') with HTTP 200.
DeepSeek-V4.1-Flash by DeepSeek costs $0.3 per 1M input tokens and $1.20 per 1M output tokens ($0.525/1M blended), with a 1M (1.000.000-token) context window. It is generally available (GA). Its weights are publicly downloadable, so it can be self-hosted on your own infrastructure.
Last verified: 26 Sep 2026 · sourced from official provider documentation
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · no spam · unsubscribe anytime.
Compare DeepSeek-V4.1-Flash head-to-head
$0.525 vs $10 blended /M · cross-provider
$0.525 vs $8 blended /M · cross-provider
$0.525 vs $1.50 blended /M · cross-provider
$0.525 vs $3 blended /M · cross-provider
$0.525 vs $0.75 blended /M · cross-provider
$0.525 vs $1.40 blended /M · cross-provider
DeepSeek-V4.1-Flash — questions & answers
How much does DeepSeek-V4.1-Flash cost?
DeepSeek-V4.1-Flash is priced at $0.3 per 1M input tokens and $1.20 per 1M output tokens ($0.525/1M blended at a 3:1 input-to-output ratio), with cached input at $0.006 per 1M tokens.
What is the context window of DeepSeek-V4.1-Flash?
DeepSeek-V4.1-Flash has a 1.000.000-token context window (1M), with up to 384.000 output tokens per request.
Is DeepSeek-V4.1-Flash deprecated?
No — DeepSeek-V4.1-Flash is generally available (GA) and not currently scheduled for deprecation or retirement in our tracker.
Does DeepSeek-V4.1-Flash support image or vision input?
Yes — DeepSeek-V4.1-Flash accepts image input. Listed modalities: text, image.
Can I self-host DeepSeek-V4.1-Flash?
Yes — DeepSeek-V4.1-Flash is an open-weight model: its weights are publicly downloadable, so you can run it on your own infrastructure (subject to its license) instead of only calling DeepSeek's API.
What is the API model string for DeepSeek-V4.1-Flash?
The API model identifier for DeepSeek-V4.1-Flash is "deepseek-flash" when calling DeepSeek's API.
Track DeepSeek-V4.1-Flash price & status changes
New models, price cuts, and deprecations — a short email when something actually changes. No spam, unsubscribe anytime.
The last few changes we caught — this is what lands in your inbox:
- DeprecatingGoogle re-pointed the Gemini 2.5 TTS migration path past the 3.1 preview and onto the new GA 3.8 TTS models
- New modelGemini 3.8 Flash TTS is GA - a flagship text-to-speech model at under half the audio rate of the preview it replaces, on a two-column price schedule
- New modelGemini 3.8 Flash-Lite TTS is GA - Google's cost-efficient TTS, which Google says was built to replace Gemini 3.1 Flash TTS
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · powered by the same data on this page.