Qwen3-Max
GAPrevious flagship Max, alias = qwen3-max-2026-01-23. Official International TIERED pricing: 0<token<=32K: $1.2 in / $6 out; 32K<token<=128K: $2.4 / $12; 128K<token<=256K: $3 / $15. Non-Thinking and Thinking modes. Context window 262,144 (256K) and max output 32,768 per OpenRouter (Alibaba spec table is SPA-rendered, not directly fetchable); knowledge cutoff Jun 2025 per OpenRouter. SCHEDULED RETIREMENT 2026-10-10 00:00:00 (UTC+08) -> replacement qwen3.7-max, per Alibaba service notice id=1950 ("Notice of Retirement for Selected Legacy Mainline Models", published 2026-06-08), which names qwen3-max, qwen3-max-preview and qwen3.6-max-preview. The dated snapshots qwen3-max-2026-01-23 and qwen3-max-2025-09-23 carry the same 2026-10-10 retirement in the companion snapshot notice id=1949. CORRECTED 2026-09-08: until today this row published 2026-09-08 as its deprecation/retirement date. That date came from the 2026-06-28 backfill and no first-party Alibaba surface has ever carried it — notice id=1950 has said 2026-10-10 since it was published on 2026-06-08, and the English policy hub (https://www.alibabacloud.com/help/en/model-studio/model-depreciation) groups this batch under "Scheduled for deprecation on October 10, 2026". Our error, not a provider change, so no changelog event. The citation moved off that policy hub because the hub no longer names any model: since its 2026-08-25 revision the per-model lists live only inside the linked service notices, and there as a table IMAGE rather than as text (path recorded so it is not re-walked). The notice's two cost columns state the REPLACEMENT's rates (qwen3.7-max), not this row's; the prices in this row remain from https://www.alibabacloud.com/help/en/model-studio/model-pricing. deprecated_on and retires_on carry the same date because Alibaba publishes exactly one: the hub calls 2026-10-10 the deprecation date and the notice calls it the retirement date.
Qwen3-Max by Alibaba costs $1.20 per 1M input tokens and $6 per 1M output tokens ($2.40/1M blended), with a 262K (262.144-token) context window. It is scheduled for retirement, shutting down 10 Oct 2026; the recommended replacement is Qwen3.7-Max.
Last verified: 26 Sep 2026 · sourced from official provider documentation
Retires on 10 Oct 2026 · deprecated since 10 Oct 2026
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · no spam · unsubscribe anytime.
This model is deprecated. Recommended migration: Qwen3.7-Max.
Source: Alibaba official documentation ↗
Qwen3-Max — questions & answers
How much does Qwen3-Max cost?
Qwen3-Max is priced at $1.20 per 1M input tokens and $6 per 1M output tokens ($2.40/1M blended at a 3:1 input-to-output ratio).
What is the context window of Qwen3-Max?
Qwen3-Max has a 262.144-token context window (262K), with up to 32.768 output tokens per request.
Is Qwen3-Max being deprecated or retired?
Qwen3-Max is scheduled for retirement. Its scheduled shutdown date is 10 Oct 2026. The recommended migration is Qwen3.7-Max.
Does Qwen3-Max support image or vision input?
Our catalog lists Qwen3-Max's modalities as text; image/vision input is not among them.
Can I self-host Qwen3-Max?
No — Qwen3-Max is proprietary. Its weights are not publicly released, so it is available only through Alibaba's API (or hosted partners), not for self-hosting.
What is the API model string for Qwen3-Max?
The API model identifier for Qwen3-Max is "qwen3-max" when calling Alibaba's API.
What is Qwen3-Max's knowledge cutoff?
Qwen3-Max's training knowledge cutoff is Jun 2025.
Track Qwen3-Max price & status changes
New models, price cuts, and deprecations — a short email when something actually changes. No spam, unsubscribe anytime.
The last few changes we caught — this is what lands in your inbox:
- DeprecatingGoogle re-pointed the Gemini 2.5 TTS migration path past the 3.1 preview and onto the new GA 3.8 TTS models
- New modelGemini 3.8 Flash TTS is GA - a flagship text-to-speech model at under half the audio rate of the preview it replaces, on a two-column price schedule
- New modelGemini 3.8 Flash-Lite TTS is GA - Google's cost-efficient TTS, which Google says was built to replace Gemini 3.1 Flash TTS
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · powered by the same data on this page.