Changelog
Every launch, price move, deprecation and retirement we track — newest first.
Google's Gemini API deprecations table changed the Recommended replacement column for both Gemini 2.5 text-to-speech endpoints. gemini-2.5-flash-preview-tts and gemini-2.5-pro-preview-tts previously pointed at gemini-3.1-flash-tts-preview (still the case on this catalog's 2026-09-24 read) and now read 'gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts' (2026-09-25 read) - the two GA models Google released on 2026-09-22, which the 3.1 preview itself was already pointing at. Google does not date the edit, so this entry carries the day it was first seen. The field records one id and carries the first-named, gemini-3.8-flash-tts; both are first-party targets. WHAT THIS IS NOT: it is not a shutdown. gemini-2.5-flash-preview-tts still reads 'No shutdown date announced' and stays preview; gemini-2.5-pro-preview-tts was already a gray (shut-down) row and stays retired - only its onward pointer moved. Against the 2.5 Flash TTS rates of $0.50 text in / $10.00 audio out per 1M, both new targets are cheaper on audio output while Google's current column runs (through 2026-12-31: $9.00 for 3.8 Flash TTS, $6.00 for 3.8 Flash-Lite TTS), and both double on 2027-01-01. Between the same two reads Google also removed a duplicate April 13, 2026 row for gemini-3.1-flash-tts-preview, leaving February 26, 2026 as its only published release date.
source ↗Google's release note for September 22, 2026 is headed 'Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS generally available (GA)' and describes gemini-3.8-flash-tts as the 'flagship creative TTS model engineered for studio-grade voice fidelity, nuanced acting, regional dialects, and long-form multi-turn stability'. The same release shipped a Voices endpoint (/v1beta/voices) with voice design, consent-verified voice replication and a 150+ voice library. Model page: 8,192-token input and 16,384-token output limits, text in, audio out, 130 languages, WAV output by default. Pricing follows Google's dated two-column shape, and this catalog's fields carry the column in force today: Standard $0.50 input (text) / $9.00 output (audio) / $0.125 cached per 1M tokens through December 31, 2026, doubling to $1.00 / $18.00 / $0.25 on January 1, 2027. Batch and Flex are half of Standard, Priority 1.8x. For comparison, gemini-3.1-flash-tts-preview carries $1.00 / $20.00. Google's deprecations table now names this model first as that preview's replacement ('gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts'), with no shutdown date for the preview. Recorded two days after Google's date: the event is theirs, the delay is ours.
source ↗Same Google release note as gemini-3.8-flash-tts, dated September 22, 2026. It describes gemini-3.8-flash-lite-tts as a 'fast, cost-efficient TTS model built to replace gemini-3.1-flash-tts-preview for high-throughput production and real-time voice agent cascades'. Model page: 8,192-token input and 16,384-token output limits, text in, audio out, 101 languages. Standard pricing per 1M tokens through December 31, 2026: $0.50 input (text) / $6.00 output (audio) / $0.125 cached, doubling to $1.00 / $12.00 / $0.25 on January 1, 2027; Batch and Flex are half of Standard, Priority 1.8x. Input matches gemini-3.8-flash-tts and output is two-thirds of it. On the lifecycle side, Google's deprecations table now names both 3.8 TTS models as replacements for gemini-3.1-flash-tts-preview, but announces no shutdown date for it, so that model stays preview; this catalog's single replacement field carries the first-named, gemini-3.8-flash-tts, and the preview's notes record both. Recorded two days after Google's date.
source ↗OpenAI's API changelog entry for September 22, 2026 reads: "Released GPT-6 Sol (gpt-6-sol) and GPT-6 Luna (gpt-6-luna). These reasoning models accept text and image inputs and generate text through the Responses and Chat Completions APIs." Standard pricing per 1M tokens for prompts up to 272K input tokens, as OpenAI states it there and on the pricing page: GPT-6 Sol $2 input / $0.20 cached input / $10 output (cache writes $2.50); GPT-6 Luna $0.10 / $0.01 / $0.50 (cache writes $0.125). Above 272K input tokens the whole request bills at 2x input and cache rates and 1.5x output - the same rule GPT-6 Astra carries. Both share Astra's 1,050,000-token context and 128,000 max output; knowledge cutoffs are April 20, 2026 (Sol, ten days before Astra's April 30) and May 18, 2026 (Luna, the latest of the three). Model pages: Sol is "built for complex coding and agentic workflows", Luna is "our most efficient model for focused, high-volume tasks". A difference from Astra that matters to anyone migrating between the three: Sol and Luna DO support reasoning.effort 'none', and Chat Completions supports function calling on them only with reasoning_effort set to none. Batch and Flex are 50% of Standard, Fast mode is 2x, and the same day the pricing page broadened its data-residency note to 'For GPT-6 Astra, Sol, and Luna, EU data residency is available only with Standard processing.' Unlike Astra's launch, neither model page carries a staged-rollout sentence, so this catalog records both as ga. OpenAI announces no deprecation of any existing model alongside this release, including gpt-5.6-luna, so no lifecycle field elsewhere moves. gpt-6-luna has its own entry for the same announcement, so a consumer filtering by model sees each. Recorded one day after OpenAI's date: the event is theirs, the day's delay is ours.
source ↗Same OpenAI changelog entry as gpt-6-sol, dated September 22, 2026: "Released GPT-6 Sol (gpt-6-sol) and GPT-6 Luna (gpt-6-luna)." Luna's Standard pricing per 1M tokens up to 272K input tokens is $0.10 input / $0.01 cached input / $0.50 output, with cache writes at $0.125; above 272K the whole request bills at 2x input and cache rates and 1.5x output. 1,050,000-token context, 128,000 max output, May 18, 2026 knowledge cutoff, text and image in, text out, reasoning.effort from none to max. OpenAI's model page calls it "our most efficient model for focused, high-volume tasks". The name overlaps with gpt-5.6-luna, but OpenAI states no relationship between the two - no deprecation of gpt-5.6-luna and no migration pointer - so that row is unchanged and none is inferred. Recorded one day after OpenAI's date: the event is theirs, the delay is ours.
source ↗xAI's release notes for September 21, 2026 state that Grok 4.7, 'SpaceXAI's frontier model for coding, agentic tasks, and knowledge work, is now available on the xAI API as grok-4.7'. Specs as xAI publishes them: 500,000-token context window, text and image input with text-only output, and an explicit 'no text output limit', which is why this catalog carries max_output_tokens as null for an absence of a limit rather than for an unknown. Pricing is tiered on prompt size: $2.00 input / $0.50 cached input / $6.00 output per 1M tokens below 200K prompt tokens, and exactly double above it ($4 / $1 / $12). THAT IS THE SAME SCHEDULE GROK 4.6 HAS CARRIED SINCE 2026-08-12, to the cent and on both tiers - a flagship generation that moved no rate at all, which is worth saying plainly because a new frontier model is usually where a price rise lands. Reasoning effort accepts low, medium, high and xhigh with high as the default, function calling and structured outputs are supported, and the Batch API is NOT - the same gap grok-4.6 has and grok-4.3 and grok-4.5 do not. New in this generation and stated in the release note: on the Responses API the model always returns reasoning.encrypted_content, even when include does not list it. The knowledge cut-off is May 2026, stated in prose on the models index with no day attached; this catalog's date field therefore carries 2026-05-01 as a convention, not as a published date. No alias is published for the id, so there is no grok-4.7-latest to pin against - the index still lists only grok-4.5-latest and grok-4.3-latest. WHAT DID NOT HAPPEN: grok-4.6 was displaced, not deprecated. xAI announces no deprecation date, no shutdown date and no migration guide for it, its model page still calls it a frontier model, and the pricing page still serves it on all three clusters at unchanged rates, so its row keeps status ga with deprecated_on, retires_on and replacement null. Also not carried here: 'Grok 4.7 Fast', which the same release note describes as the same model at twice the token rates and states is available only through Cursor and Grok Build, not on the public xAI API - a model this catalog cannot price against a public rate card is one it does not list.
source ↗DeepSeek has revised the footnote that defines its peak window on the Models & Pricing page. It now reads: 'Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday, excluding Chinese public holidays. All other hours are off-peak, including weekends and Chinese public holidays in full.' The previous wording, which this catalog quoted in its 2026-08-14, 2026-08-16 and 2026-09-10 entries and in both live DeepSeek rows, was: 'Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak).' NOT A SINGLE RATE CHANGED - this is a change to WHEN the half rate applies, not to what it is. Both live models keep their published pair: deepseek-flash at peak $0.006 cache-hit / $0.30 cache-miss / $1.20 output per 1M, deepseek-v4-pro at peak $0.044 / $1.32 / $3.96, off-peak exactly half of each. What a consumer gains: on a Chinese public holiday that falls on a weekday, the 7 hours that used to bill at peak now bill at half, so an evenly-spread workload that runs right through one of those days pays about 23% less than the old schedule would have charged it (24 hours at the half rate = 24 half-rate-hours, against 7 peak + 17 off-peak = 31 under the old wording). The carve-out cuts both ways as a planning problem - DeepSeek publishes no calendar of which days those are, and the Chinese public-holiday calendar is set annually by the State Council, so the billing schedule now depends on a list that is not on this page. THIS ENTRY IS DATED BY OBSERVATION, NOT BY THE PROVIDER. DeepSeek published no Change Log entry for the revision and put no date on the footnote; it edited the page in place, exactly as it did when it withdrew the V4 Pro retirement on 2026-09-12. The only bound is this catalog's own daily prose baseline of that page, which carried the old sentence on 2026-09-18 and the new one on 2026-09-19. Separately, and visible in the same read: the footnote that carried the V4 Pro reprieve has been removed from the pricing page now that September 14 is past. That is housekeeping, not a reversal - the identical paragraph is still in the Change Log entry for 2026-09-10, V4 Pro is still in the rate table on its own unchanged rates, and its row keeps deprecated_on, retires_on and replacement null.
source ↗xAI released grok-voice-transcribe-2.0 on 2026-09-17 and dated it itself, in its release notes: 'grok-voice-transcribe-2.0 is now available. Use grok-voice-transcribe-1.0 or grok-voice-transcribe-2.0.' THE PRICE DID NOT MOVE AND THE OLD MODEL DID NOT GO AWAY - which is the whole story here, because a version 2 usually means one or the other. xAI's pricing page publishes a single Speech to Text line covering the mode rather than a line per model, at $0.10 per hour of audio for the REST path and $0.20 per hour for streaming, and the models-index payload carries byte-identical rate integers for both slugs, so a caller upgrading pays exactly what they paid before and a caller staying put is not on a deprecated endpoint: xAI announces no deprecation, no shutdown date and no successor for 1.0, and this catalog therefore leaves all three lifecycle fields null rather than inferring a retirement from the arrival of a 2. ONE THING TO SET EXPLICITLY: the default is contradicted across two first-party surfaces. The release note says 'the default is grok-voice-transcribe-1.0'; the Speech to Text reference prints grok-voice-transcribe-2.0 in the model parameter's default cell. Pass model yourself and the question does not arise. Capabilities on the 2.0 endpoint: 25 languages, batch and streaming, 12 audio formats, word-level timestamps, multichannel transcription across 2-8 channels, speaker diarization, Smart Turn end-of-turn detection, a tunable vad_threshold, 500 MB per file, 10 rps / 600 rpm / 100 concurrent sessions at the base tier. THIS ENTRY ARRIVED TWO DAYS LATE: the event and its date are xAI's, the delay in recording it is ours. It landed in the same read that added xAI's other three audio models to this catalog, none of which had ever been carried.
source ↗Google released gemini-3.8-live on 2026-09-15 and dated it itself: the Release date column of the deprecations table reads "September 15, 2026", and all three Google surfaces - the models index, the model page and the pricing page - are stamped "Last updated 2026-09-15 UTC". It arrives as STABLE, not preview: the models index files it under "Gemini 3 / Stable" with a "New" badge and the model page's Versions row reads "Stable: gemini-3.8-live". Google calls it the "Default Live API model for most low-latency voice agent experiences without reasoning delays". THE PRICE STORY IS THAT THERE ISN'T ONE. Google did not give the new model its own rate card - it rewrote the pricing page's Live section to head a SINGLE Standard table with "Gemini 3.8 Live, Gemini 3.8 Live Extended Thinking, and Gemini 3.1 Flash Live Preview", so the figures the catalog already carried for the 3.1 preview now apply to all three ids unchanged: input $0.75 (text) / $3.00 or $0.005 per minute (audio) / $1.00 or $0.002 per minute (image/video); output $4.50 (text) / $12.00 or $0.018 per minute (audio). Anyone migrating off gemini-3.1-flash-live-preview pays exactly what they paid before. Caching is not supported and no cached-input rate is published. Capacity: 131,072 input / 65,536 output tokens, text/images/audio/video in, text and audio out. The migration is not a drop-in string swap, and Google says so on the model page under "Migrating from Gemini 3.1 Flash Live": thinking_level is not supported and must be omitted, asynchronous function calling (behavior: NON_BLOCKING) becomes the default mode, proactive audio is permanently on and setting proactive_audio: false now returns an error, and affective dialogue has been removed from the API entirely, so any enable_affective_dialog configuration has to come out of the code. Google announces no shutdown date for this model and the deprecations row is not gray.
source ↗Google released gemini-3.8-live-extended-thinking on 2026-09-15, the same day as gemini-3.8-live and on the same evidence: "September 15, 2026" in the Release date column of the deprecations table, "Stable: gemini-3.8-live-extended-thinking" in the model page's Versions row, and a "New - Stable" badge on the models index. Google calls it the "High-reasoning Live API model for voice interactions, recommended when higher background reasoning is required", and says it "processes background reasoning and asynchronous tool calls while streaming continuous audio responses". It is priced identically to gemini-3.8-live on every axis, because Google publishes ONE Standard table for both of them plus gemini-3.1-flash-live-preview: input $0.75 (text) / $3.00 or $0.005 per minute (audio) / $1.00 or $0.002 per minute (image/video); output $4.50 (text) / $12.00 or $0.018 per minute (audio). The extra reasoning carries no rate premium in anything Google publishes today. Capacity matches too: 131,072 input / 65,536 output tokens, text/images/audio/video in, text and audio out. ONE capability line separates the two rows, and it is the one that costs work: function calling here is "Supported (Async only)", where gemini-3.8-live still accepts synchronous blocking calls for backwards compatibility. Note also what this model is NOT: Google does not name it as the replacement for any endpoint. All four Live endpoints it re-pointed the same day point at gemini-3.8-live. No shutdown date is announced and the deprecations row is not gray.
source ↗On 2026-09-15 Google changed the Recommended replacement column of its deprecations table for four Live API endpoints in one edit. Three of them - gemini-2.0-flash-live-001, gemini-live-2.5-flash-preview and gemini-2.5-flash-native-audio-preview-12-2025 - previously pointed at gemini-3.1-flash-live-preview and now point at gemini-3.8-live. The fourth is gemini-3.1-flash-live-preview itself, which had no replacement at all and now has one: the model Google had been sending everyone TO is now the model it is sending people OFF. The models index states it in plain words, describing the 3.1 preview as a "Legacy Live API preview model. We recommend updating to Gemini 3.8 Live." WHAT THIS IS NOT: it is not a shutdown and not a date. Google still states "No shutdown date announced" for gemini-3.1-flash-live-preview and for gemini-2.5-flash-native-audio-preview-12-2025, and neither row is gray, so both keep their preview status here - a re-pointed migration target is a statement about where to go, not about when the old endpoint stops answering. The two rows that ARE gray (gemini-2.0-flash-live-001 and gemini-live-2.5-flash-preview, both shut down 2025-12-09) were already retired and stay retired; only their onward pointer moved. This event class is invisible to a price diff - no rate moved, and in fact the 3.1 preview's rates are now published in the same table as the model replacing it, so a consumer watching only prices would see a completely flat day.
source ↗DeepSeek has withdrawn the retirement of DeepSeek-V4-Pro that it announced on 2026-09-10 and that was due to take effect at 12:00 Beijing Time on 2026-09-14. The model stays on the API with its rates untouched. DeepSeek states it as footnote (2) beside the live rate column on its Models & Pricing page, and again in the Change Log entry where the retirement text used to sit: 'In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged. We will provide further notice should there be any changes. Thank you for your understanding and support!' Read the last clause as part of the announcement: this is a withdrawal, not a commitment, and DeepSeek reserves the right to re-announce. THIS ENTRY IS DATED BY OBSERVATION, NOT BY THE PROVIDER. DeepSeek published no new Change Log entry for the reversal - it edited its 2026-09-10 entry in place, so that entry now reads as though the retirement was never announced at all. The only bound on the publication date is this catalog's own daily prose baseline of the pricing page, which carried the retirement paragraph on 2026-09-11 and the withdrawal paragraph on 2026-09-12; the true date is one of those two. Anyone who read the 2026-09-10 entry when it was published, or the deprecation notice this catalog issued from it, saw a retirement that is no longer scheduled. What did NOT change: V4 Pro is still priced at peak $0.044 cache-hit / $1.32 cache-miss / $3.96 out per 1M with off-peak exactly half, still 1M context and 384K max output, and still without vision support. What is now unknown again: DeepSeek named no successor and no date, and the V4.1 Pro it referred to as unreleased on 2026-09-10 is still unreleased, so this row carries no replacement and no shutdown date. The two endpoints retired alongside it on 2026-09-10, deepseek-v4-flash and deepseek-v4-flash-vision-exp, were NOT reinstated: the pricing page still says those names are accepted but the models behind them are retired and requests are served by DeepSeek-V4.1-Flash.
source ↗DeepSeek released DeepSeek-V4.1-Flash on 2026-09-10 and dated it itself, in its API Change Log: "Today, we officially release the DeepSeek-V4.1-Flash model. It is the smallest model in our new architecture family, with native multimodal visual understanding. The new architecture is designed for a higher capability ceiling, faster inference, higher throughput, and scaling to larger models." Call it as deepseek-flash - that is the model name DeepSeek instructs you to use, and it is not the same string as this catalog's row id. Prices are a PEAK / OFF-PEAK pair, and the fields here carry PEAK: $0.30 per 1M cache-miss input, $1.20 per 1M output, $0.006 per 1M cache-hit input, with off-peak exactly half on every axis ($0.15 / $0.60 / $0.003). DeepSeek defined the peak window narrowly - "Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak)" - so off-peak is what most callers pay most of the time. (That wording is the one DeepSeek published on this date; it revised the footnote on or just before 2026-09-19 to exclude Chinese public holidays from peak and make them off-peak in full. The rates in this entry are unaffected - see the 2026-09-19 entry.) Against the model it replaces, deepseek-v4-flash at $0.44 / $1.32 / $0.014 peak, that is roughly 32% off input, 9% off output and 57% off cache-hit input, and DeepSeek states the reduction as such: "With the release of DeepSeek-V4.1-Flash, API prices have been reduced accordingly." The headline capability change is vision: the feature table marks Vision supported here and "Not supported" for deepseek-v4-pro, and the weights are tagged image-text-to-text. Everything else carries over - 1M context, 384K max output, thinking and non-thinking modes with thinking as the default, JSON output, tool calls, Responses API, Anthropic-format API, chat prefix completion, FIM completion in non-thinking mode only, concurrency limit 2500. Open weights went up the same day at huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash under the MIT licence. This catalog carries it as GA rather than preview because DeepSeek announces it as an official release with no preview or experimental label, unlike the V4 generation it shipped as "DeepSeek V4 Preview". Knowledge cutoff is not published, so that field is null.
source ↗On 2026-09-10 DeepSeek retired two endpoints in one sentence of its API Change Log: "The previous-generation models V4 Flash and V4 Flash Vision Exp have been retired; for compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily routed to V4.1 Flash." The pricing page footnote says the same thing and adds the billing consequence: the legacy names "are still accepted, but the corresponding models have been retired, their requests are served by the DeepSeek-V4.1-Flash model and billed at the Flash price." That is an explicit, past-tense, dated statement by the provider - the standard this catalog requires for a retirement - and not an inference from the columns disappearing off the pricing page, which they also did the same day. Two things matter for anyone consuming this. First, nothing breaks: calls to either legacy name still succeed, so a service left unmigrated will not fail loudly, it will silently start getting a different model, at the new and lower Flash price. DeepSeek calls that routing "temporary" and names no end date for it, so the migration to deepseek-flash is real work with no deadline attached. Second, prices on both retired rows are frozen at the rates in force at retirement - peak $0.014 cache-hit / $0.44 cache-miss / $1.32 out for each, off-peak half - and deliberately no longer track the endpoint now serving the traffic. DeepSeek published no separate shutdown date and no deprecation window for either, so each row carries the single date it did publish in retires_on and leaves deprecated_on null. deepseek-v4-flash-vision-exp enters this catalog on the day it leaves the API: an experimental vision model announced 2026-08-21 and retired twenty days later, added here so the lifecycle record is complete. Its rates were recovered from an Internet Archive snapshot of DeepSeek's own pricing page taken 2026-09-07, because the live page had already dropped the column by the time the retirement was caught.
source ↗WITHDRAWN - READ THIS FIRST. DeepSeek reversed this announcement before it took effect: on or about 2026-09-12 it replaced the retirement paragraph, on both the pricing page and this very Change Log entry, with 'In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged.' V4 Pro was never retired and carries no shutdown date today. This entry is kept unedited below because it is an accurate record of what DeepSeek published on 2026-09-10 and of what this catalog served for two days; DeepSeek itself did not keep such a record, having edited the announcement away in place. See the 2026-09-12 entry for the withdrawal. WHAT WAS ANNOUNCED ON 2026-09-10, verbatim as it stood: DeepSeek announced on 2026-09-10 that it will retire DeepSeek-V4-Pro - its flagship, the ~1.6T-parameter top of the V4 line - four days later. Its Change Log states the whole thing: "extensive testing shows that V4.1 Flash now outperforms DeepSeek V4 Pro across performance, cost, speed, and total time, so we plan to retire V4 Pro in an orderly manner. After 12:00 Beijing Time on September 14, 2026, and until the future release of V4.1 Pro, all requests to deepseek-v4-pro will be routed to V4.1 Flash and billed at the V4.1 Flash price." Read the time, not only the date: 12:00 Beijing is 04:00 UTC on 2026-09-14, and a boundary finer than a day matters to anyone billing against it. Until then V4 Pro is unchanged and still priced at peak $0.044 cache-hit / $1.32 cache-miss / $3.96 out, off-peak exactly half. After it, the same model name returns a smaller model at roughly a quarter of the cost - $0.006 / $0.30 / $1.20 at peak - and, notably, a model that supports vision where V4 Pro does not. There is no Pro-tier successor to migrate to: DeepSeek says V4.1 Pro is unreleased and gives no date, so the replacement field on this row points at deepseek-v4-1-flash, which is where the traffic actually goes, rather than at a model that does not yet exist. As with the two endpoints retired the same day, calls will not fail - which is exactly what makes this worth watching, since the failure mode is a silent substitution of the model behind a stable name rather than an error a monitor would catch.
source ↗GPT Image 2.5 Sunburst is described as OpenAI's most capable model for image generation and editing, for workflows where editing precision matters most. OpenAI's API changelog for Sep 8 reads: 'Released GPT Image 2.5 Sunburst and GPT Image 2.5 Flare for image generation and editing through the Image API and the Responses API image generation tool.' Both support the new xhigh and max quality settings. Rates are GPT Image 2's: $8 image input (cached $2) and $30 image output per 1M tokens, plus $5 text input (cached $1.25) per 1M. OpenAI's pricing page now shows the two 2.5 models first and collapses gpt-image-2 among the older image models; gpt-image-2 is not deprecated. Recorded 2026-09-26, 18 days after OpenAI's date: the event is theirs, the delay is ours.
source ↗GPT Image 2.5 Flare is described as OpenAI's fastest model for high-quality, everyday image generation. OpenAI's API changelog for Sep 8 reads: 'Released GPT Image 2.5 Sunburst and GPT Image 2.5 Flare for image generation and editing through the Image API and the Responses API image generation tool.' Both support the new xhigh and max quality settings. Rates are GPT Image 2's: $8 image input (cached $2) and $30 image output per 1M tokens, plus $5 text input (cached $1.25) per 1M. OpenAI's pricing page now shows the two 2.5 models first and collapses gpt-image-2 among the older image models; gpt-image-2 is not deprecated. Recorded 2026-09-26, 18 days after OpenAI's date: the event is theirs, the delay is ours.
source ↗Google re-priced both Robotics-ER 2 endpoints on its Gemini API pricing page, stamped "Last updated 2026-09-08 UTC", and the new cells are a DATED TWO-COLUMN SCHEDULE rather than a single number. For gemini-robotics-er-2-preview, paid-tier standard per 1M: input $1.00 (text / image / video / audio) through December 31, 2026 then $2.00 starting January 1, 2027; output including thinking tokens $5.00 through December 31, 2026 then $10.00 starting January 1, 2027; context caching $0.10 through December 31, 2026 then $0.20 starting January 1, 2027, plus a storage fee of $0.50 per 1M tokens per hour through December 31, 2026 then $1.00 starting January 1, 2027. Batch moves in step at half of standard: $0.50/$2.50 through December 31, 2026, then $1.00/$5.00. gemini-robotics-er-2-streaming-preview gets the same standard schedule ($1.00 in / $5.00 out today, $2.00/$10.00 from January 1, 2027) and still publishes no context-caching rate and no batch tier at all. Read that as a promotional window, not a permanent reduction: the January 1, 2027 column is exactly the flat $2.00/$10.00/$0.20 Google had charged since the 2026-07-30 launch, so the effect is a 50% discount for the rest of 2026 that expires back to the launch price. Two things worth knowing about how this is recorded here. First, the price fields in this catalog carry the column IN FORCE TODAY, and each row's notes state that explicitly and spell out the whole schedule — a schedule collapsed to one number is a silent 2x error the day it turns over, which is the failure mode that had gemini-3.6-flash serving Google's 2027 rate as current for about two weeks. Second, nobody has to remember the cutover: check-google-prices.mjs parses the schedule and resolves it against the clock, so these cells move on their own from its scheduled-change list to a hard DRIFT report on 2027-01-01. The change is Google's, not a correction of ours: the rates published here until today were first-party-sourced and were what Google charged for the period they covered, and the 2026-07-30 launch entry in this changelog records them as such.
source ↗Google's Gemini API deprecations page now renders the veo-2.0-generate-001 row with a grey background, and that page's own legend reads "Already-shutdown models are indicated with gray backgrounds" - a first-party, per-row, machine-readable assertion that the endpoint is off, carried as a literal class="row-gray" on the <tr>. check-lifecycle-drift.mjs reads that marker and reported Google stating "shut-down" against this catalog's "deprecated" on 2026-09-05. The row was ungreyed on the 2026-09-04 read, so the marker appeared between the two. Google publishes no separate execution date; the shutdown date it had published - June 30, 2026 - passed 67 days earlier without the endpoint going off, and that gap is exactly why this catalog does not retire on a calendar: Google states its shutdown dates as "the earliest possible dates on which a model might be retired". Veo 2 was a text-to-video model billed per second of video at $0.35/s, paid tier only, with no token pricing. Migration target, from the same table: veo-3.1-generate-preview, or the GA models on the Gemini Enterprise Agent Platform. One caveat worth publishing: a retirement for this row filed on 2026-08-31 was reverted the same day, because it rested on the rate row vanishing from the pricing page - an absence, not a statement. Today's signal is the grey row itself.
source ↗Google's Gemini API release notes for September 3, 2026 read "Lyria 3.5 in public preview: Released the next generation of Google's music generation model: lyria-3.5: Full-length song generation with improved musical coherence, natural vocals, and fine-grained duration and structural control." It takes text AND image inputs through the Interactions API - Lyria 3 is text-in only - and emits 44.1 kHz high-fidelity stereo audio, MP3 by default with WAV available via response_format. Pricing is per request and identical to the full-song tier it succeeds: $0.08 per song, paid tier only with no free tier, so the per-token fields are N/A rather than unknown. The second half of this is the part a consumer would otherwise miss: the same week, the pricing page re-headed the Lyria 3 block "Google's family of legacy music generation models" and the deprecations table began naming lyria-3.5 as the recommended replacement for lyria-3-pro-preview. That pointer has been adopted here. Google publishes NO replacement for lyria-3-clip-preview and NO deprecation date and NO shutdown date for either Lyria 3 endpoint, so both keep status preview: a replacement target is not a lifecycle date, and this catalog does not infer one. ENTRY CORRECTED 2026-09-05. As first published it named two ids, lyria-3.5-clip-preview and lyria-3.5-pro-preview, read from a pricing-page listing that stood for about a day. Google's own announcement named one model, and every Google surface names only lyria-3.5 today - release notes, pricing page, music-generation guide and deprecations table. The two rows were withdrawn as our error, not retired; no shutdown was ever declared for them.
source ↗OpenAI's API changelog entry for September 3, 2026 reads "Released GPT-6 Astra, our most capable model, built for the hardest end-to-end work", tagged v1/responses and v1/chat/completions. Specs from its model page: 1,050,000-token context, 128,000 max output, April 30 2026 knowledge cutoff, text+image input and text output, reasoning token support. Standard tier per 1M: $10.00 input / $1.00 cached input / $12.50 cache write / $50.00 output; Batch and Flex are 50% of Standard, Fast mode is 2x and is unavailable with EU data residency; prompts over 272K input tokens bill at 2x input and cache rates and 1.5x output for the whole request. Migration notes OpenAI states itself: reasoning.effort has no 'none' level on this model, custom temperature/top_p and logprobs are unsupported, and tool calling requires the Responses API. This catalog records it as status "preview" and that is a judgement call worth reading before depending on it: OpenAI never calls it preview, but it also never says "generally available" — a phrase it uses freely elsewhere in the same changelog — and the model page carries the sentence "GPT-6 Astra is rolling out today for enterprises in our Trusted Access Program, with access through API and our Plus, Pro, Business and Enterprise plans coming in the coming days." Broad API access is therefore announced as pending, so "ga" would tell a consumer they can call it today. Rate limits are nevertheless already published for usage tiers 1-5.
source ↗Google released gemini-3.8-flash on September 2, 2026, under a release note headed "Gemini 3.8 Flash generally available (GA)" and describing it as "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows". Specs from its model page: 1,048,576-token input / 65,536-token output, inputs text/image/video/audio/PDF, text output, thinking at low/medium/high ('minimal' returns an error), caching, code execution, function calling, file search, search and Maps grounding, URL context and computer use (preview); no Live API, no audio generation, no image generation. Stable channel is the bare id. The price fields carry the column IN FORCE TODAY: $0.75 in / $3.75 out / $0.075 cached per 1M through December 31, 2026, doubling to $1.50 / $7.50 / $0.15 on January 1, 2027 — the third Gemini Flash row on this dated-schedule shape, which is the class of error that had gemini-3.6-flash publishing the 2027 column as today's rate for two weeks. Google does NOT use the word 'introductory' for 3.8 anywhere it uses it for 3.7, so this catalog does not either. Batch and Flex are $0.375 / $1.875 (caching $0.0375), Priority $1.35 / $6.75 (caching $0.135), each doubling on the same date. Same launch rewrote the pricing page's one-line descriptions across the whole Flash ladder — 3.7 Flash lost the 'most capable Flash model' billing to 3.8 — but no field of any existing row changed, so that is recorded in the affected rows' notes and not as an event here.
source ↗xAI opened a 60-day notice period on September 2, 2026 for the grok-imagine-image-quality slug, stating on its migration guide that 'Effective November 2, 2026, the grok-imagine-image-quality model slug is retired from the xAI API' and that 'Its 60-day notice period began on September 2, 2026'. THE UNUSUAL PART IS THAT NOTHING BREAKS: after the date, requests to the retired slug on /v1/images/generations and /v1/images/edits are served by grok-imagine-image-2.0 with quality set to 'low', request and response shapes unchanged, and the response's model field reports which model actually served the request. xAI states the redirect target is $0.01 per image cheaper at every resolution than the model it replaces, so callers who do nothing pay less, not more - the reason to migrate explicitly is to choose a quality tier instead of inheriting the 'low' default. grok-imagine-image-pro, retired 2026-05-15 and already redirecting to grok-imagine-image-quality, follows the chain on to grok-imagine-image-2.0. grok-imagine-image (1.0) is explicitly not affected. This catalog carried the row as 'ga' with no lifecycle dates for the 16 days between the announcement and 2026-09-18; the delay is ours, the event and both its dates are xAI's.
source ↗Moonshot's model list (platform.kimi.ai/docs/models.md) carries a warning block reading "The kimi-k2.5 and moonshot-v1 series were officially retired on August 31, 2026. Calls to these models now return a 404 (model not found) error. Please migrate to kimi-k3." That is a first-party statement of an executed shutdown with an observable consequence, not a scheduled date — the strongest class of evidence this catalog accepts, and the reason this row moved on the read rather than on the calendar. Corroborated the same morning by Moonshot deleting the model's own pricing page: platform.kimi.ai/docs/pricing/chat-k25.md now 307-redirects onto the model list. Kimi K2.5 was a 1T-parameter MoE (~32B active) native multimodal model with a 262,144-token context, priced $0.60 cache-miss input / $0.10 cache-hit / $3.00 output per 1M. Migration target: kimi-k3 ($3.00 / $0.30 / $15.00 per 1M, 1,048,576-token context), which this catalog did not carry until today. Read the correction honestly: this row published status "ga" for two days after Moonshot switched the endpoint off, because no instrument in this repo reads Moonshot — check-lifecycle-drift covers Google, Anthropic and Mistral, and check-pricing-prose covers five providers that do not include it. The kimi-k2 series (0711, 0905, turbo, thinking, thinking-turbo) is NOT part of this retirement: Moonshot still lists those as "Deprecated" and the 404 warning does not name them.
source ↗Moonshot retired the entire moonshot-v1 family on August 31, 2026: "The moonshot-v1 series models (including moonshot-v1-auto and the -vision-preview variants) were officially discontinued on August 31, 2026", and the page's warning block adds that calls to them "now return a 404 (model not found) error". Six of the seven ids are carried here and all six changed status today: moonshot-v1-8k, moonshot-v1-32k and moonshot-v1-128k were published as "ga"; moonshot-v1-8k-vision-preview, moonshot-v1-32k-vision-preview and moonshot-v1-128k-vision-preview as "preview". The seventh, moonshot-v1-auto, has never been in this catalog and is not being added as a retired row. Corroborating first-party signal: platform.kimi.ai/docs/pricing/chat-v1.md, which is where every one of these rows was priced, now 307-redirects onto the model list — Moonshot deleted the price page along with the models, so the rates carried here are explicitly the last published ones (moonshot-v1-8k at $0.20 in / $2.00 out per 1M, read 2026-06-28) and not a live quote. Migration target for all six: kimi-k3.
source ↗OpenAI's API changelog for August 21, 2026 states: "GPT-5.6 Sol now costs $4 per million input tokens and $20 per million output tokens, representing 20% lower input pricing and 33% lower output pricing. GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026." The model page repeats both percentages and its pricing block now reads $4.00 input / $0.40 cached input / $20.00 output; cache writes are $5.00/M at 1.25x uncached input. Batch and Flex remain 50% of Standard, Fast mode 2x. This is filed on the provider's date, not the date it was read: the rows here carried $5.00/$0.50/$30.00 until 2026-09-04, so the feed served rates 25% high on input and 50% high on output for fourteen days. Cause, recorded because it is the actionable part: scripts/check-price-drift.mjs reads exactly this page and exactly these rows, and it is not in the daily data routine's step list — an instrument that exists in the repo but not in the playbook does not run. The rate is explicitly promotional with a published floor of November 21, 2026, not a confirmed durable price.
source ↗Executed. DeepSeek's price rise, announced with a date and a table on 2026-08-14, takes effect at 16:00 UTC today, and both V4 rows have been re-pointed from the old flat rate to the new PEAK column. deepseek-v4-flash moves from $0.0028 cache-hit / $0.14 cache-miss / $0.28 output per 1M to $0.014 / $0.44 / $1.32; deepseek-v4-pro moves from $0.003625 / $0.435 / $0.87 to $0.044 / $1.32 / $3.96. That is 4.71x and 4.55x on output respectively, and 12.1x on pro's cache-hit input — the largest single multiple in the table. Two things a consumer of this feed should know, because the schema has one price field and DeepSeek now publishes two. First, we carry PEAK because DeepSeek presents peak as the price and off-peak as half of it, not the reverse — so the fields are the ceiling, and the OFF-PEAK rate is exactly half on every axis. Second, peak is the minority of the day: 01:00-04:00 and 06:00-10:00 UTC, 7 hours of 24, so a workload that does not care when it runs pays the half rate for 17 hours a day. The legacy deepseek-chat and deepseek-reasoner aliases are NOT affected and keep their old figures: they were retired 2026-07-24 15:59 UTC, three weeks before this schedule began, so it never applied to them.
source ↗DeepSeek's pricing page has replaced the undated warning it carried since at least 2026-08-01 — 'We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected' — with a dated schedule. From 16:00 UTC on August 16, 2026 the API moves to peak/off-peak billing, with off-peak rates stated as half the peak rates; peak hours are 01:00-04:00 and 06:00-10:00 UTC, so 7 of every 24 hours are peak and the rest are off-peak. Per 1M tokens, deepseek-v4-flash goes from $0.0028 cache-hit / $0.14 cache-miss / $0.28 output to OFF-PEAK $0.007 / $0.22 / $0.66 and PEAK $0.014 / $0.44 / $1.32 — 4.71x on peak output and 2.36x even off-peak. deepseek-v4-pro goes from $0.003625 / $0.435 / $0.87 to OFF-PEAK $0.022 / $0.66 / $1.98 and PEAK $0.044 / $1.32 / $3.96 — 4.55x on peak output, and 12.1x on peak cache-hit input, the largest single multiple in the table. This entry is the ANNOUNCEMENT; the executed re-point is the separate 2026-08-16 entry, where the catalog fields moved to the PEAK column, since DeepSeek presents peak as the price and off-peak as half of it. Caught by scripts/check-pricing-prose.mjs, which diffs a provider’s own sentences rather than its tables — the change was invisible to every rate diff we run, because on 2026-08-14 not one published number had moved yet.
source ↗Google's release notes for August 13, 2026 make Gemini 3.7 Flash (gemini-3.7-flash) generally available, calling it 'our most intelligent workhorse model yet for coding and agents' with substantial improvements across software engineering, web development and agentic workflows. Specs: 1,048,576-token context, 65,536 max output, text/image/video/audio/PDF in and text out, thinking at low/medium/high ('minimal' returns an error), batch, flex and priority inference all supported. Google publishes no knowledge cutoff, and the deprecations table announces no shutdown date. The pricing is the part worth reading carefully: paid-tier Standard is $0.75 in / $3.75 out per 1M with context caching at $0.075/1M — but only 'through December 31, 2026'. From January 1, 2027 the same page states $1.50 / $7.50 / $0.15, exactly double on all three axes, and Google's own release note calls the current rate an introductory price. That makes 3.7 Flash identical in price to gemini-3.6-flash today and identical again after the cutover. Anyone budgeting off today's number is budgeting off a rate with a published expiry date; our fields carry the in-force column and a dated ticket moves them on the boundary.
source ↗xAI's release notes for August 12, 2026 state that Grok 4.6, 'SpaceXAI's frontier model for coding, agentic tasks, and knowledge work, is now available on the xAI API'. Specs as xAI publishes them: 500,000-token context window, text and image input with text-only output, and - unusually, because most providers simply omit the field - an explicit 'no text output limit', which is why this catalog carries max_output_tokens as null for an absence of a limit rather than for an unknown. Pricing is tiered on prompt size: $2.00 input / $0.50 cached input / $6.00 output per 1M tokens below 200K prompt tokens, and exactly double above it ($4 / $1 / $12). Reasoning effort accepts low, medium, high and xhigh with high as the default, function calling and structured outputs are supported, and the Batch API is NOT - both grok-4.3 and grok-4.5 accept batch requests and 4.6 does not. The knowledge cut-off is February 1, 2026, stated in prose on the models index. No alias is published for the id, so there is no grok-4.6-latest to pin against. It displaces grok-4.5 as the model the models index recommends for general use ('For everything else, including code, use Grok 4.6'), without any lifecycle change to 4.5 or 4.3: xAI announces no deprecation or shutdown date for either. This catalog did not carry the model at all for the 37 days between the launch and 2026-09-18; the delay is ours, the launch and its date are xAI's.
source ↗Anthropic's deprecations table now states Current state = Retired for claude-opus-4-1-20250805, with Deprecated June 5, 2026 and Retirement August 5, 2026 — exactly the date announced two months earlier. Calls to claude-opus-4-1-20250805 (and the claude-opus-4-1 alias) no longer serve; Anthropic's recommended replacement is claude-opus-4-8. Opus 4.1 was $15 in / $75 out per 1M tokens, 200K context, 32K max output. Caught by scripts/check-lifecycle-drift.mjs on 2026-08-06 — the provider's declared status flipped while our catalog still read 'deprecated'. Worth noting for anyone building on rate data: this event is invisible to a price diff, because the price never changed — the endpoint simply stopped existing.
source ↗Google's deprecations table now names gemini-3.6-flash — not gemini-3.5-flash — as the recommended replacement for gemini-2.5-flash (shutdown 2026-10-16), gemini-2.5-flash-preview-05-20, gemini-2.5-flash-preview-09-25, and the whole gemini-2.0-flash line (gemini-2.0-flash and gemini-2.0-flash-001, both shut down 2026-06-01). The new target is CHEAPER on output than the old one: Gemini 3.6 Flash is $0.75 in / $3.75 out per 1M (context caching $0.075) against Gemini 3.5 Flash's $1.50 / $9.00, so the two Flash rows are not a straight ladder. Gemini 3.6 Flash is GA/stable, 1,048,576-token context, 65,536 max output, text+image+video+audio+PDF input; Google publishes no knowledge cutoff for it. CORRECTED 2026-08-15: this entry originally quoted Gemini 3.6 Flash at $1.50 / $7.50 / $0.15 and said the two models shared an input price. Those are Google's 'starting January 1, 2027' figures — the page prices this model on a two-date schedule and the rates above are the ones in force through 2026-12-31, so 3.6 Flash is currently cheaper than 3.5 Flash on BOTH axes and the gap closes on 2027-01-01. models.json was corrected on 2026-08-14; this entry, which had copied the same wrong column, was missed for a day. The release date this entry said Google did not publish is in the deprecations table: July 21, 2026.
source ↗OpenAI's API changelog for July 30, 2026 states: "Starting July 30, GPT-5.6 Luna costs 80% less, while GPT-5.6 Terra costs 20% less." That reconciles exactly with the launch rates published here on 2026-07-09 — Luna $1.00 input / $0.10 cached / $6.00 output falls to $0.20 / $0.02 / $1.20, and Terra $2.50 / $0.25 / $15.00 falls to $2.00 / $0.20 / $12.00 — and both current figures were verified against the pricing page and each model page again on 2026-09-04. This event is being filed retroactively on 2026-09-04, and the reason matters more than the numbers: on 2026-08-08 a run noticed both rows disagreed with the pricing page, changed the fields to the correct values, and wrote them up in each row's notes as "PRICE CORRECTED — we published X, which appears in no cell of OpenAI's pricing page". It was not our arithmetic; it was a dated provider price cut, which under this project's rules is a changelog event with an alert, while a correction of our own is neither. Nine days of drift were read as a typo, so no event was recorded and no consumer of the feed was told the price had moved. The same July 30 entry also renamed Priority processing to Fast mode (2x Standard, up to 2.5x faster on Sol), which is a tier change and not a rate change to any row here.
source ↗Google released two new embodied-reasoning endpoints for robotics in public preview: gemini-robotics-er-2-preview (advanced spatial reasoning, agentic code execution, multi-step tool orchestration, video moment finding, progress classification, multi-robot coordination) and gemini-robotics-er-2-streaming-preview (real-time text streaming over the Live API for low-latency robot agents with bidirectional audio and video input). Both accept text, image, video and audio input, output text, and carry 131,072-token input / 65,536-token output limits. Paid-tier standard rates are $2.00/M in and $10.00/M out for both — 2x the input and 2x the output of Gemini Robotics-ER 1.6 ($1.00/$5.00). Only the non-streaming endpoint publishes a context-caching rate ($0.20/1M) and a batch tier ($1.00/$5.00). Google's own upgrade guidance is to replace model="gemini-robotics-er-1.6-preview" with either ER 2 endpoint; ER 1.6 is shut down 2026-08-31.
source ↗grok-voice-think-fast-2.0, xAI's realtime speech-to-speech model for voice agents, was released on 2026-07-29 and dated by xAI's own release notes: 'grok-voice-think-fast-2.0 is now available with Speech to Speech. grok-voice-latest will route to this model starting August 5, 2026.' That second sentence is the part worth acting on: anyone pinned to the alias was moved onto this model on 2026-08-05 without changing a string, which is exactly the class of change this catalog exists to make visible - an alias is a price pointer and a behaviour pointer, not only a version pointer. It is driven over a WebSocket at wss://api.x.ai/v1/realtime, with low-latency turn-taking and server-side tool use, and it is billed by connected audio time rather than by token: $0.08 per minute of audio, which xAI also prints as $4.80/hr, plus $0.01 per minute for sessions bridged to the public telephone network. Concurrency, not throughput, is the limit axis - 10 simultaneous sessions at the base tier, rising to 400 at tier 5. THIS ENTRY IS SEVEN WEEKS LATE AND THE DELAY IS OURS, NOT A PROVIDER EVENT: the date, the alias cutover and the rates are all xAI's own published figures, and this catalog simply never carried the model. It was added on 2026-09-19 together with grok-tts and both Grok Voice Transcribe endpoints, closing the audio line that a read on 2026-09-18 had flagged as the one known gap left at this provider. Two further rate fields in xAI's payload are not carried at all, because neither first-party surface says what unit they are in; the model row explains which and why.
source ↗DeepSeek announced, unconditionally and with a timestamp, that "deepseek-chat & deepseek-reasoner will be fully retired and inaccessible after Jul 24th, 2026, 15:59 (UTC Time)". Both rows move from deprecated to retired. The evidence class is worth stating plainly because it is weaker than the OpenAI entry above and stronger than date arithmetic alone: DeepSeek publishes no live status table to re-read, so there is no past-tense confirmation to find; what there is, is a declaration with no "earliest possible date" hedge, plus absence — neither id appears on api-docs.deepseek.com/quick_start/pricing/ nor in the current list-models documentation, both re-read 2026-09-02. Absence from a pricing page is not sufficient on its own (it is exactly the signal that wrongly retired six Google rows on 2026-08-31 and had to be reverted); here it corroborates a first-party declaration rather than substituting for one. These were aliases, not models: deepseek-chat routed to deepseek-v4-flash non-thinking and deepseek-reasoner to its thinking mode. Their price fields remain the rates in force at retirement and deliberately do not track deepseek-v4-flash, which moved to peak/off-peak billing on 2026-08-16, three weeks after these aliases stopped answering.
source ↗OpenAI's deprecations page has rewritten the April 2026 batch in the past tense: the section is now headed "2026-04-22: Legacy GPT model snapshots (July 2026 shutdown)" and states "Access to these models was shut down on July 23, 2026." Three of the shut-down ids are carried here and all three move from deprecated to retired: gpt-5-codex, o3-deep-research and computer-use-preview. This is a deliberate distinction, not a formality — the shutdown date alone had already passed on all three and was not treated as evidence, because a published date is a plan and a past-tense sentence is a fact. The substitutes changed too, and that is the part worth acting on: OpenAI now names gpt-5.6-sol for gpt-5-codex (was gpt-5.5 here) and for o3-deep-research (was gpt-5.5-pro), and gpt-5.6-terra for computer-use-preview (was gpt-5.4-mini). All three replacement pointers were re-pointed in the same edit. Note for anyone reading the same page: the identically dated "2026-04-22: Legacy GPT model snapshots" section without the suffix is a DIFFERENT batch, still scheduled, shutting down October 23, 2026.
source ↗Google shipped stable, production-ready Gemini 3.6 Flash and Gemini 3.5 Flash-Lite on the same day. Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite) is billed at $0.30/M in and $2.50/M out on the paid tier, context caching $0.03/M, batch $0.15/$1.25 — against gemini-3.1-flash-lite's $0.25/$1.50, i.e. the recommended migration target is more expensive on both sides, unusually for a Flash-Lite step. One offsetting simplification: 3.5 Flash-Lite publishes a SINGLE input rate covering text, image, video and audio, where 3.1 and 2.5 Flash-Lite both charged a higher audio-input tier. 1,048,576-token context, 65,536 max output, text/image/video/audio/PDF in, text out; computer use in preview. Google's deprecations table names it the recommended replacement for gemini-3.1-flash-lite, which shuts down 2027-05-07. No shutdown date announced for 3.5 Flash-Lite itself.
source ↗OpenAI's GPT-5.6 generation went GA on the API + Codex on 2026-07-09 (preview from 06-26), in three durable capability tiers: Sol (flagship) $5 in / $30 out, Terra (balanced) $2.50 / $15, Luna (fast, high-volume) $1 / $6 — all per 1M tokens, cached input $0.50 / $0.25 / $0.10. Every tier has a 1.05M-token context window, 128K max output, text + image input, and a 2026-02-16 knowledge cutoff. Sol matches GPT-5.5's price but is stronger on coding, knowledge work, cybersecurity and science. New caching model: explicit cache breakpoints, 30-min minimum cache life, cache writes billed 1.25x uncached input. [Appended 2026-09-04: these are the LAUNCH rates and three of them have since been cut by OpenAI — Terra -20% and Luna -80% on 2026-07-30, Sol to $4 in / $20 out on 2026-08-21. See the dated price entries for those days; do not read this entry as current pricing.]
source ↗xAI's smartest model for chat, coding, agentic and knowledge work. Base pricing $2 input / $6 output per 1M tokens (cached input $0.50), tiered 2x above a 200K-token prompt. 500K context window (down from Grok 4.3's 1M), text + image input. Public API access opened 2026-07-09; aliases grok-4.5-latest / grok-build-latest.
source ↗Anthropic's most agentic Sonnet, with performance close to Opus 4.8 at lower cost. Standard pricing $3 input / $15 output per 1M tokens, with introductory pricing of $2/$10 in effect through 2026-08-31. 1M context window, 128k max output, Jan 2026 knowledge cutoff. Supersedes Claude Sonnet 4.6, now a legacy model.
source ↗Mistral released Leanstral 1.5, a specialist model for Lean 4 formal proof engineering, automated theorem proving and autoformalization (not a general chat model). Sparse MoE 119B total / 6.5B active, 256K context, 128K max output, text + image input. Open weights under Apache 2.0. Offered FREE in Mistral's experimental Labs tier ($0/M in and out), with a stated retirement of 2026-09-30 — the catalog's first $0 model, recorded as preview (Labs/experimental) so it stays out of the GA 'cheapest' rankings.
source ↗Migrate to amazon.nova-2-lite-v1:0. Where that shutdown stands, checked first-hand on 2026-09-15, the day after the stated date: AWS's model card still reads 'Model lifecycle: Legacy' and still carries 'Model EOL date: September 14, 2026', so AWS has not declared the shutdown executed and this catalog keeps the row at deprecated rather than retired. A date that has merely passed is not a retirement signal. The same card also states 'EOL no sooner than: Oct 31, 2026' beside that EOL date — two first-party fields that contradict each other; both are recorded, neither is resolved into a typed field.
source ↗Migrate to qwen3.7-max. ENTRY CORRECTED 2026-09-08: this entry published "shuts down 2026-09-08". Alibaba service notice id=1950 ("Notice of Retirement for Selected Legacy Mainline Models"), published 2026-06-08 — three weeks before this entry was written — states 2026-10-10 00:00:00 (UTC+08), and no first-party Alibaba surface has ever carried 2026-09-08. Our error, not a provider change. The citation moved to the notice because the policy hub stopped naming individual models in its 2026-08-25 revision.
source ↗Migrate to qwen3.7-max. ENTRY CORRECTED 2026-09-08: this entry published "shuts down 2026-09-08". Alibaba service notice id=1950 ("Notice of Retirement for Selected Legacy Mainline Models"), published 2026-06-08 — three weeks before this entry was written — states 2026-10-10 00:00:00 (UTC+08), and no first-party Alibaba surface has ever carried 2026-09-08. Our error, not a provider change. The citation moved to the notice because the policy hub stopped naming individual models in its 2026-08-25 revision.
source ↗Migrate to qwen3.7-plus. Citation repointed 2026-09-10: this entry cited Alibaba's model decommissioning policy page, which was rewritten on 2026-08-25 to drop its per-model tables and delegate to individual service notices, so it no longer names this model or any other. The certifying document is service notice id=1949, 'Notice of Retirement for Selected Legacy Snapshot Models' (published 2026-06-08), whose table names qwen3-235b-a22b alongside qwen3-235b-a22b-instruct-2507 and -thinking-2507 and states the retirement as 2026-10-10 00:00:00 (UTC+08). The model row itself was corrected on 2026-09-08; this entry was left on the dead citation until today, which is the recurring pattern that a correction reaches the row it was about and not the other surfaces publishing the same fact.
source ↗Tracking 181 models across 12 providers.
OpenAI's deprecations page, under '2026-06-02: GPT Image model deprecations': 'On June 2, 2026, we notified developers using older GPT Image models of their deprecation and removal from the API on December 1, 2026.' The table lists gpt-image-1-mini, gpt-image-1.5 and the chatgpt-image-latest alias, each with gpt-image-2 as the recommended replacement. Note that gpt-image-2 has since been joined by GPT Image 2.5 Sunburst and Flare (2026-09-08) at the same token rates, but the table still names gpt-image-2. THIS ENTRY IS ALMOST FOUR MONTHS LATE AND THE DELAY IS OURS: the date, the shutdown and the replacement are OpenAI's own published figures; this catalog carried the model as ga until 2026-09-26.
source ↗OpenAI's deprecations page, under '2026-06-02: GPT Image model deprecations': 'On June 2, 2026, we notified developers using older GPT Image models of their deprecation and removal from the API on December 1, 2026.' The table lists gpt-image-1-mini, gpt-image-1.5 and the chatgpt-image-latest alias, each with gpt-image-2 as the recommended replacement. Note that gpt-image-2 has since been joined by GPT Image 2.5 Sunburst and Flare (2026-09-08) at the same token rates, but the table still names gpt-image-2. THIS ENTRY IS ALMOST FOUR MONTHS LATE AND THE DELAY IS OURS: the date, the shutdown and the replacement are OpenAI's own published figures; this catalog carried the model as ga until 2026-09-26.
source ↗Mistral deprecated its frontier reasoning model (magistral-medium-2509, $2/$5 per MTok, 128k context). Migrate to Mistral Medium 3.5 — Mistral is folding its specialist reasoning line back into its general-purpose models rather than shipping a Magistral successor.
source ↗Mistral deprecated its frontier agentic-coding model (devstral-2512, $0.4/$2 per MTok, 256k context, open weights under a Modified MIT licence). Migrate to Mistral Medium 3.5 — the specialist coding line is being folded back into the general-purpose models.
source ↗Mistral deprecated its small reasoning model (magistral-small-2509, $0.5/$1.5 per MTok, 128k context, Apache-2.0 weights, 24B). Migrate to Mistral Small 4.
source ↗The snapshot o4-mini-2025-04-16 and the o4-mini alias that resolves to it both stop serving on 2026-10-23. OpenAI names gpt-5.6-terra as the replacement. Until the shutdown the model is still sold at the Standard rates of $1.10 input, $0.275 cached input and $4.40 output per 1M tokens, so a caller sees no price signal that anything is ending. A fine-tuned build of the same snapshot, ft-o4-mini-2025-04-16, carries the same date in a separate table. OpenAI announced this on 2026-04-22 under the heading '2026-04-22: Legacy GPT model snapshots', which states the shutdown date and the recommended replacement as columns of a table: "To improve reliability and make it easier for developers to choose the right models, we are deprecating a set of older OpenAI models. Access to these models will be shut down on the dates below." THIS ENTRY IS FIVE MONTHS LATE AND THE DELAY IS OURS: the date, the shutdown and the replacement are all OpenAI's own published figures, and this catalog had never carried the model at all. That same announcement has a sibling section, "(July 2026 shutdown)", whose models this catalog did record on the day - it carried one half of one announcement and missed the other, and the half it missed is the one still live.
source ↗The snapshot o3-mini-2025-01-31 and the o3-mini alias stop serving on 2026-10-23, with gpt-5.6-sol named as the replacement. Standard rates until then are $1.10 input, $0.55 cached input and $4.40 output per 1M. Worth noting for anyone choosing between the two small reasoning models in their last month: o3-mini and o4-mini are priced identically on input and output and differ only in cached input, where o3-mini costs twice as much. OpenAI announced this on 2026-04-22 under the heading '2026-04-22: Legacy GPT model snapshots', which states the shutdown date and the recommended replacement as columns of a table: "To improve reliability and make it easier for developers to choose the right models, we are deprecating a set of older OpenAI models. Access to these models will be shut down on the dates below." THIS ENTRY IS FIVE MONTHS LATE AND THE DELAY IS OURS: the date, the shutdown and the replacement are all OpenAI's own published figures, and this catalog had never carried the model at all. That same announcement has a sibling section, "(July 2026 shutdown)", whose models this catalog did record on the day - it carried one half of one announcement and missed the other, and the half it missed is the one still live.
source ↗The snapshot o1-pro-2025-03-19 and the o1-pro alias stop serving on 2026-10-23. The replacement cell is the unusual one in this group: OpenAI writes gpt-5.6-sol and then qualifies it with "reasoning.mode: pro", i.e. the migration is a parameter on another model, not a swap of model string. Standard rates until then are $150.00 input and $600.00 output per 1M with no cached-input rate published, which makes this by a wide margin the most expensive row OpenAI still sells. OpenAI announced this on 2026-04-22 under the heading '2026-04-22: Legacy GPT model snapshots', which states the shutdown date and the recommended replacement as columns of a table: "To improve reliability and make it easier for developers to choose the right models, we are deprecating a set of older OpenAI models. Access to these models will be shut down on the dates below." THIS ENTRY IS FIVE MONTHS LATE AND THE DELAY IS OURS: the date, the shutdown and the replacement are all OpenAI's own published figures, and this catalog had never carried the model at all. That same announcement has a sibling section, "(July 2026 shutdown)", whose models this catalog did record on the day - it carried one half of one announcement and missed the other, and the half it missed is the one still live.
source ↗The snapshot o1-2024-12-17 and the o1 alias stop serving on 2026-10-23, with gpt-5.6-sol named as the replacement. Standard rates until then are $15.00 input, $7.50 cached input and $60.00 output per 1M. The cached-input discount on o1 is 50%, against the 90% OpenAI applies across the gpt-5.x line, so a cache-heavy workload is paying a larger premium to stay here than the headline input rate shows. o1-preview and o1-mini were separate, earlier deprecations whose shutdown dates (2025-07-28 and 2025-10-27) have already passed. OpenAI announced this on 2026-04-22 under the heading '2026-04-22: Legacy GPT model snapshots', which states the shutdown date and the recommended replacement as columns of a table: "To improve reliability and make it easier for developers to choose the right models, we are deprecating a set of older OpenAI models. Access to these models will be shut down on the dates below." THIS ENTRY IS FIVE MONTHS LATE AND THE DELAY IS OURS: the date, the shutdown and the replacement are all OpenAI's own published figures, and this catalog had never carried the model at all. That same announcement has a sibling section, "(July 2026 shutdown)", whose models this catalog did record on the day - it carried one half of one announcement and missed the other, and the half it missed is the one still live.
source ↗The gpt-4.1-nano alias and its gpt-4.1-nano-2025-04-14 snapshot stop serving on 2026-10-23, with gpt-5.6-luna named as the replacement. Standard rates until then are $0.10 input, $0.025 cached input and $0.40 output per 1M. THE MIGRATION IS NOT PRICE-NEUTRAL: gpt-5.6-luna is priced at $0.20 input and $1.20 output per 1M on the same table, so following the recommendation doubles the input rate and triples the output rate. Its cached-input rate is slightly lower in absolute terms, so a workload that is almost entirely cache hits is the one case that does not get more expensive. OpenAI announced this on 2026-04-22 under the heading '2026-04-22: Legacy GPT model snapshots', which states the shutdown date and the recommended replacement as columns of a table: "To improve reliability and make it easier for developers to choose the right models, we are deprecating a set of older OpenAI models. Access to these models will be shut down on the dates below." THIS ENTRY IS FIVE MONTHS LATE AND THE DELAY IS OURS: the date, the shutdown and the replacement are all OpenAI's own published figures, and this catalog had never carried the model at all. That same announcement has a sibling section, "(July 2026 shutdown)", whose models this catalog did record on the day - it carried one half of one announcement and missed the other, and the half it missed is the one still live.
source ↗OpenAI lists gpt-4o-2024-05-13 alone in this group, not the gpt-4o alias, so what ends on 2026-10-23 is the May 2024 snapshot and anyone on the bare alias is not affected by this row. gpt-5.6-sol is named as the replacement. The snapshot is still sold at $5.00 input and $15.00 output per 1M with no cached-input rate published, which is twice the input rate the live alias carries on the same table - pinning to the original snapshot has been costing double for some time, and now it also has an end date. OpenAI announced this on 2026-04-22 under the heading '2026-04-22: Legacy GPT model snapshots', which states the shutdown date and the recommended replacement as columns of a table: "To improve reliability and make it easier for developers to choose the right models, we are deprecating a set of older OpenAI models. Access to these models will be shut down on the dates below." THIS ENTRY IS FIVE MONTHS LATE AND THE DELAY IS OURS: the date, the shutdown and the replacement are all OpenAI's own published figures, and this catalog had never carried the model at all. That same announcement has a sibling section, "(July 2026 shutdown)", whose models this catalog did record on the day - it carried one half of one announcement and missed the other, and the half it missed is the one still live.
source ↗gpt-4-turbo, the gpt-4-turbo-2024-04-09 snapshot and the gpt-4-turbo-completions alias all stop serving on 2026-10-23, with gpt-5.6-sol named as the replacement. Standard rates until then, listed against the dated snapshot, are $10.00 input and $30.00 output per 1M with no cached-input rate published. The separate gpt-4-turbo-preview line, which resolves to gpt-4-0125-preview, was an earlier deprecation on the same page with a shutdown date of 2026-03-26 that has already passed. OpenAI announced this on 2026-04-22 under the heading '2026-04-22: Legacy GPT model snapshots', which states the shutdown date and the recommended replacement as columns of a table: "To improve reliability and make it easier for developers to choose the right models, we are deprecating a set of older OpenAI models. Access to these models will be shut down on the dates below." THIS ENTRY IS FIVE MONTHS LATE AND THE DELAY IS OURS: the date, the shutdown and the replacement are all OpenAI's own published figures, and this catalog had never carried the model at all. That same announcement has a sibling section, "(July 2026 shutdown)", whose models this catalog did record on the day - it carried one half of one announcement and missed the other, and the half it missed is the one still live.
source ↗OpenAI lists gpt-4-0613 with the gpt-4, gpt-4-0613-completions and gpt-4-completions aliases folded into the same row, all stopping on 2026-10-23, replaced by gpt-5.6-sol. This is the end of the original GPT-4 chat endpoint on the API. Standard rates until then are $30.00 input and $60.00 output per 1M with no cached-input rate published, i.e. it is still priced above every model OpenAI recommends in its place. A fine-tuned build, ft-gpt-4, carries the same date in a separate table. OpenAI announced this on 2026-04-22 under the heading '2026-04-22: Legacy GPT model snapshots', which states the shutdown date and the recommended replacement as columns of a table: "To improve reliability and make it easier for developers to choose the right models, we are deprecating a set of older OpenAI models. Access to these models will be shut down on the dates below." THIS ENTRY IS FIVE MONTHS LATE AND THE DELAY IS OURS: the date, the shutdown and the replacement are all OpenAI's own published figures, and this catalog had never carried the model at all. That same announcement has a sibling section, "(July 2026 shutdown)", whose models this catalog did record on the day - it carried one half of one announcement and missed the other, and the half it missed is the one still live.
source ↗gpt-3.5-turbo-0125, together with the gpt-3.5-turbo alias and gpt-3.5-turbo-completions, stops serving on 2026-10-23, with gpt-5.6-terra named as the replacement. Standard rates until then are $0.50 input and $1.50 output per 1M, with no cached-input rate published and the alias and the dated snapshot priced identically. The separately priced gpt-3.5-turbo-1106 and gpt-3.5-turbo-instruct rows are NOT in this deprecation group and OpenAI states no end date for them. A fine-tuned build, ft-gpt-3.5-turbo, carries the same date in a separate table. OpenAI announced this on 2026-04-22 under the heading '2026-04-22: Legacy GPT model snapshots', which states the shutdown date and the recommended replacement as columns of a table: "To improve reliability and make it easier for developers to choose the right models, we are deprecating a set of older OpenAI models. Access to these models will be shut down on the dates below." THIS ENTRY IS FIVE MONTHS LATE AND THE DELAY IS OURS: the date, the shutdown and the replacement are all OpenAI's own published figures, and this catalog had never carried the model at all. That same announcement has a sibling section, "(July 2026 shutdown)", whose models this catalog did record on the day - it carried one half of one announcement and missed the other, and the half it missed is the one still live.
source ↗Video generation (Videos API). Priced per second, not per token: sora-2 $0.10/s (720p, standard) / $0.05/s batch. Videos API + sora-2* announced for discontinua
source ↗Higher-quality video model. Per-second pricing: $0.30/s (720p) up to $0.70/s (1080p) standard; batch half. Discontinued with Videos API, shuts down 2026-09-24,
source ↗Mistral deprecated its small agentic-coding model (labs-devstral-small-2512, $0.1/$0.3 per MTok, 256k context, Apache-2.0 weights, 24B). Migrate to Mistral Medium 3.5. Its stated 2026-03-31 retirement has passed, but Mistral still lists the model as deprecated rather than retired and still prices it publicly.
source ↗No replacement MODEL — Cohere directs users to the Chat endpoint with a summarization prompt, which is an endpoint change, not a model swap.
source ↗Get the next change before your code breaks
New models, price cuts, and deprecations — a short email when something actually changes. No spam, unsubscribe anytime.
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · powered by the same data on this page.