Kimi K2.5
RetiredNative multimodal model (text, image, video input), 1T-param MoE (~32B active). Release date 2026-01-27 from multiple secondary sources (Baidu Baike, ComfyUI Wiki, AWS Bedrock card, kimi.com). Knowledge cutoff and max output tokens never published. Confirmed RETIRED on 2026-08-31. Signal (first-party, stated as a consequence, not inferred): platform.kimi.ai/docs/models.md carries a warning block reading "The kimi-k2.5 and moonshot-v1 series were officially retired on August 31, 2026. Calls to these models now return a 404 (model not found) error. Please migrate to kimi-k3" — read 2026-09-02. Corroborated the same morning by Moonshot deleting this model's own pricing page: platform.kimi.ai/docs/pricing/chat-k25.md and chat-v1.md now 307-redirect onto the model list, which is why the citation below was repointed there. Until 2026-09-02 this row published status "ga", which was wrong from 2026-08-31 onward — nothing in this repo reads Moonshot, so the retirement went unseen for two days. Prices retained as the last published rates (cache-miss input $0.60/M, cache-hit $0.10/M, output $3.00/M), read from platform.kimi.ai/docs/pricing/chat-k25.md on 2026-06-28 before Moonshot deleted that page. Moonshot names kimi-k3 as the migration target.
Kimi K2.5 by Moonshot costs $0.6 per 1M input tokens and $3 per 1M output tokens ($1.20/1M blended), with a 262K (262.144-token) context window. It has been retired (shut down 31 Aug 2026); the recommended migration is Kimi K3. Its weights are publicly downloadable, so it can be self-hosted on your own infrastructure.
Last verified: 26 Sep 2026 · sourced from official provider documentation
Retired on 31 Aug 2026
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · no spam · unsubscribe anytime.
This model has been retired. Recommended migration: Kimi K3.
Read the Kimi K2.5 migration guide → replacement & cheapest alternatives
Kimi K2.5 — questions & answers
How much does Kimi K2.5 cost?
Kimi K2.5 is priced at $0.6 per 1M input tokens and $3 per 1M output tokens ($1.20/1M blended at a 3:1 input-to-output ratio), with cached input at $0.1 per 1M tokens.
What is the context window of Kimi K2.5?
Kimi K2.5 has a 262.144-token context window (262K).
Is Kimi K2.5 still available?
Kimi K2.5 has been retired and is no longer available. Its scheduled shutdown date is 31 Aug 2026. The recommended migration is Kimi K3.
Does Kimi K2.5 support image or vision input?
Yes — Kimi K2.5 accepts image input. Listed modalities: text, image, video.
Can I self-host Kimi K2.5?
Yes — Kimi K2.5 is an open-weight model: its weights are publicly downloadable, so you can run it on your own infrastructure (subject to its license) instead of only calling Moonshot's API.
What is the API model string for Kimi K2.5?
The API model identifier for Kimi K2.5 is "kimi-k2.5" when calling Moonshot's API.
Track Kimi K2.5 price & status changes
New models, price cuts, and deprecations — a short email when something actually changes. No spam, unsubscribe anytime.
The last few changes we caught — this is what lands in your inbox:
- DeprecatingGoogle re-pointed the Gemini 2.5 TTS migration path past the 3.1 preview and onto the new GA 3.8 TTS models
- New modelGemini 3.8 Flash TTS is GA - a flagship text-to-speech model at under half the audio rate of the preview it replaces, on a two-column price schedule
- New modelGemini 3.8 Flash-Lite TTS is GA - Google's cost-efficient TTS, which Google says was built to replace Gemini 3.1 Flash TTS
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · powered by the same data on this page.