DeepSeek-V4-Flash-Vision-Exp
RetiredExperimental multimodal vision-understanding sibling of DeepSeek-V4-Flash, and one of the shortest-lived endpoints in this catalog: released 2026-08-21 and retired 2026-09-10, twenty days later. Both dates are DeepSeek's own, from its Change Log — 'Today, the new multimodal vision understanding model DeepSeek-V4-Flash-Vision-Exp is now available on the DeepSeek API platform. This is an experimental model that can be accessed by setting model=deepseek-v4-flash-vision-exp' (2026-08-21), and 'The previous-generation models V4 Flash and V4 Flash Vision Exp have been retired; for compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily routed to V4.1 Flash' (2026-09-10). RETIREMENT SIGNAL: an explicit past-tense statement by the provider, on a dated announcement — not an inference from the model disappearing off the pricing page, though it did also disappear from it the same day. The name is still ACCEPTED by the API, so calls do not fail; they are served by DeepSeek-V4.1-Flash and billed at the Flash price, and DeepSeek calls that routing 'temporary'. On DeepSeek's own benchmarks it was on par with V4-Flash on pure text and a significant step up on visual agent tasks. PRICES ABOVE ARE THE RATES IN FORCE AT RETIREMENT and deliberately stop tracking anything: PEAK per 1M was $0.014 cache-hit / $0.44 cache-miss / $1.32 out, off-peak exactly half ($0.007 / $0.22 / $0.66), identical to deepseek-v4-flash on every axis, with images converted to input tokens by dimension. Those figures come from DeepSeek's pricing page as it stood on 2026-09-07, read from the Internet Archive snapshot cited below, because the live page dropped the column on the day of retirement — this row was added on 2026-09-10, after the fact, so the archived first-party document is the only surface that still carries them. Context 1M, max output 384K, concurrency limit 2500, JSON output / tool calls / Responses API / Anthropic API / chat prefix completion supported, FIM completion NOT supported (it was the only V4 endpoint where FIM was unavailable). Knowledge cutoff never published. Open weights at huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp.
DeepSeek-V4-Flash-Vision-Exp by DeepSeek costs $0.44 per 1M input tokens and $1.32 per 1M output tokens ($0.66/1M blended), with a 1M (1.000.000-token) context window. It has been retired (shut down 10 Sep 2026); the recommended migration is DeepSeek-V4.1-Flash. Its weights are publicly downloadable, so it can be self-hosted on your own infrastructure.
Last verified: 26 Sep 2026 · sourced from official provider documentation
Retired on 10 Sep 2026
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · no spam · unsubscribe anytime.
This model has been retired. Recommended migration: DeepSeek-V4.1-Flash.
Read the DeepSeek-V4-Flash-Vision-Exp migration guide → replacement & cheapest alternatives
DeepSeek-V4-Flash-Vision-Exp — questions & answers
How much does DeepSeek-V4-Flash-Vision-Exp cost?
DeepSeek-V4-Flash-Vision-Exp is priced at $0.44 per 1M input tokens and $1.32 per 1M output tokens ($0.66/1M blended at a 3:1 input-to-output ratio), with cached input at $0.014 per 1M tokens.
What is the context window of DeepSeek-V4-Flash-Vision-Exp?
DeepSeek-V4-Flash-Vision-Exp has a 1.000.000-token context window (1M), with up to 384.000 output tokens per request.
Is DeepSeek-V4-Flash-Vision-Exp still available?
DeepSeek-V4-Flash-Vision-Exp has been retired and is no longer available. Its scheduled shutdown date is 10 Sep 2026. The recommended migration is DeepSeek-V4.1-Flash.
Does DeepSeek-V4-Flash-Vision-Exp support image or vision input?
Yes — DeepSeek-V4-Flash-Vision-Exp accepts image input. Listed modalities: text, image.
Can I self-host DeepSeek-V4-Flash-Vision-Exp?
Yes — DeepSeek-V4-Flash-Vision-Exp is an open-weight model: its weights are publicly downloadable, so you can run it on your own infrastructure (subject to its license) instead of only calling DeepSeek's API.
What is the API model string for DeepSeek-V4-Flash-Vision-Exp?
The API model identifier for DeepSeek-V4-Flash-Vision-Exp is "deepseek-v4-flash-vision-exp" when calling DeepSeek's API.
Track DeepSeek-V4-Flash-Vision-Exp price & status changes
New models, price cuts, and deprecations — a short email when something actually changes. No spam, unsubscribe anytime.
The last few changes we caught — this is what lands in your inbox:
- DeprecatingGoogle re-pointed the Gemini 2.5 TTS migration path past the 3.1 preview and onto the new GA 3.8 TTS models
- New modelGemini 3.8 Flash TTS is GA - a flagship text-to-speech model at under half the audio rate of the preview it replaces, on a two-column price schedule
- New modelGemini 3.8 Flash-Lite TTS is GA - Google's cost-efficient TTS, which Google says was built to replace Gemini 3.1 Flash TTS
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · powered by the same data on this page.