Migrate off Summarize (legacy endpoint)
RetiredSummarize (legacy endpoint) has been retired and shut down on 15 Sep 2025. Cohere has not named an official replacement. The cheapest comparable model we track is Llama 3.1 8B Instruct (Meta) at $0.022 per 1M tokens blended.
Last verified: 26 Sep 2026 · sourced from official provider documentation
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · no spam · unsubscribe anytime.
Budget-friendly replacements for Summarize (legacy endpoint)
Generally-available models that match Summarize (legacy endpoint)'s capabilities, ranked by blended price. Every number is computed from official catalog prices.
| Model | Blended /1M | Input / output | Context |
|---|---|---|---|
| Llama 3.1 8B Instruct Meta | $0.022 | $0.02 / $0.03 | 128K |
| Amazon Nova Micro Amazon | $0.061 | $0.035 / $0.14 | 128K |
| Command R7B Cohere | $0.066 | $0.037 / $0.15 | 128K |
Migrating off Summarize (legacy endpoint)
What replaces Summarize (legacy endpoint)?
Cohere has not published an official replacement for Summarize (legacy endpoint). The cheapest comparable model in our catalog is Llama 3.1 8B Instruct (Meta) at $0.022/1M blended.
Is Summarize (legacy endpoint) still available?
No — Summarize (legacy endpoint) has been retired (shut down 15 Sep 2025) and is no longer available via Cohere's API.
What is the cheapest alternative to Summarize (legacy endpoint)?
Llama 3.1 8B Instruct (Meta) is the cheapest generally-available alternative we track that matches Summarize (legacy endpoint)'s capabilities, at $0.02 per 1M input and $0.03 per 1M output tokens ($0.022/1M blended).
Browse the full deprecation watch → or see Summarize (legacy endpoint)'s full spec sheet →
Get a heads-up before your model is retired
New models, price cuts, and deprecations — a short email when something actually changes. No spam, unsubscribe anytime.
The last few changes we caught — this is what lands in your inbox:
- DeprecatingGoogle re-pointed the Gemini 2.5 TTS migration path past the 3.1 preview and onto the new GA 3.8 TTS models
- New modelGemini 3.8 Flash TTS is GA - a flagship text-to-speech model at under half the audio rate of the preview it replaces, on a two-column price schedule
- New modelGemini 3.8 Flash-Lite TTS is GA - Google's cost-efficient TTS, which Google says was built to replace Gemini 3.1 Flash TTS
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · powered by the same data on this page.