Qwen-Flash

GA

Legacy stable Flash alias = qwen-flash-2025-07-28; the recommended replacement for the discontinued Qwen-Turbo. Official International TIERED pricing: 0<token<=256K $0.05 in / $0.4 out; 256K<token<=1M $0.25 in / $2 out per 1M. Supports 50% batch-inference discount and context caching. Context window 1M. Max output not on accessible official page (null). Still GA on pricing page Jun 22, 2026. Text-only.

Qwen-Flash by Alibaba costs $0.05 per 1M input tokens and $0.4 per 1M output tokens ($0.138/1M blended), with a 1M (1.000.000-token) context window. It is generally available (GA).

Last verified: 13 Aug 2026 · sourced from official provider documentation

Provider
Alibaba
Status
GA
Input price
$0.05 / 1M tokens
Output price
$0.4 / 1M tokens
Cached input
Blended price
$0.138 / 1M tokens
Context window
1.000.000 tokens (1M)
Max output
Modality
text
Open weights
No — API only
Knowledge cutoff
Released
28 Jul 2025
API string
qwen-flash

Source: Alibaba official documentation ↗

FAQ

Qwen-Flash — questions & answers

How much does Qwen-Flash cost?

Qwen-Flash is priced at $0.05 per 1M input tokens and $0.4 per 1M output tokens ($0.138/1M blended at a 3:1 input-to-output ratio).

What is the context window of Qwen-Flash?

Qwen-Flash has a 1.000.000-token context window (1M).

Is Qwen-Flash deprecated?

No — Qwen-Flash is generally available (GA) and not currently scheduled for deprecation or retirement in our tracker.

Does Qwen-Flash support image or vision input?

Our catalog lists Qwen-Flash's modalities as text; image/vision input is not among them.

Can I self-host Qwen-Flash?

No — Qwen-Flash is proprietary. Its weights are not publicly released, so it is available only through Alibaba's API (or hosted partners), not for self-hosting.

What is the API model string for Qwen-Flash?

The API model identifier for Qwen-Flash is "qwen-flash" when calling Alibaba's API.