Between July 1 and September 30, 2026, Anthropic, OpenAI, Google, DeepSeek and xAI made at least 28 changes to the prices of their text-model APIs, counting new models and their launch prices. Most new flagships arrived at or below the price of the model they replaced, one vendor raised its prices, and several prices now carry an end date. Every row below is dated from the vendor’s own changelog or pricing page.
- 01Launches held the lineOpus 5, Sonnet 5.5 and Grok 4.6 and 4.7 matched their predecessors. Opus 5.5 came in 20% below Opus 5.
- 02Caching got cheaperFable 5.1, Opus 5.5 and GPT-6.1 Sol all cut the price of reading cached input.
- 03Prices with expiry datesGemini 3.6, 3.7 and 3.8 Flash double on January 1; GPT-5.6 Sol’s promotion runs at least to November 21.
- 04Time of day now countsSince August 16, DeepSeek bills peak and off-peak rates, with off-peak at half the peak rate. Both sit above its old flat price.
01 — The dataThe quarter’s price changes, dated
Prices are US dollars per million input and output tokens at the standard tier, for prompts under any long-context threshold. Where a vendor’s current page no longer shows an old price, it was read from an Internet Archive copy of that vendor’s own pricing page.
| Date | Model | Price | Change |
|---|---|---|---|
| Jul 8 | xAI Grok 4.5 | $2 / $6 | New model |
| Jul 9 | OpenAI GPT-5.6 Sol | $5 / $30 | New model at GPT-5.5’s price |
| Jul 9 | OpenAI GPT-5.6 Terra | $2.50 / $15 | New model |
| Jul 9 | OpenAI GPT-5.6 Luna | $1 / $6 | New model |
| Jul 21 | Google Gemini 3.6 Flash | $1.50 / $7.50 | Generally available; same input price as 3.5 Flash ($1.50 / $9), lower output |
| Jul 21 | Google Gemini 3.5 Flash-Lite | $0.30 / $2.50 | Generally available |
| Jul 24 | Anthropic Claude Opus 5 | $5 / $25 | New model at the same price as Opus 4.8 |
| Jul 30 | OpenAI GPT-5.6 Luna | $0.20 / $1.20 | Price cut by 80%, from $1 / $6 |
| Jul 30 | OpenAI GPT-5.6 Terra | $2 / $12 | Price cut by 20%, from $2.50 / $15 |
| Jul 30 | OpenAI Fast mode | 2× standard | Replaces Priority Processing; for GPT-5.6 Sol it costs twice standard |
| Aug 7 | OpenAI GPT-5.6 Cyber | $12.50 / $75 | New model, approved Daybreak program access only |
| Aug 10 | Anthropic Claude Sonnet 5 | $2 / $10 | Introductory price made standard; the Sep 1 rise to $3 / $15 cancelled |
| Aug 12 | xAI Grok 4.6 | $2 / $6 | New model at Grok 4.5’s price |
| Aug 13 | Google Gemini 3.7 Flash | $0.75 / $3.75 | New model at an introductory price to Dec 31; $1.50 / $7.50 from Jan 1, 2027 |
| Aug 13 | Google Gemini 3.6 Flash | $0.75 / $3.75 | Cut to an introductory price to Dec 31; back to $1.50 / $7.50 from Jan 1, 2027 |
| Aug 16 | DeepSeek API | $0.44 / $1.32 peak | Peak and off-peak billing replaces a flat rate; V4 Flash rises from $0.14 / $0.28 to $0.44 / $1.32 at peak, off-peak is half |
| Aug 21 | OpenAI GPT-5.6 Sol | $4 / $20 | Promotional price, input 20% and output 33% lower, at least to Nov 21 |
| Sep 1 | Anthropic Claude Fable 5.1 | $10 / $50 | New model; cache reads $0.25, down from $1 on Fable 5 |
| Sep 2 | Google Gemini 3.8 Flash | $0.75 / $3.75 | New model at an introductory price to Dec 31; $1.50 / $7.50 from Jan 1, 2027 |
| Sep 3 | OpenAI GPT-6 Astra | $10 / $50 | New model; $20 / $75 for prompts above 272K tokens |
| Sep 8 | OpenAI GPT-Rosalind | $5 / $25 | New model, trusted access only; billing starts Oct 5 |
| Sep 10 | DeepSeek V4.1 Flash | $0.30 / $1.20 | Replaces V4 Flash ($0.44 / $1.32 peak); peak rates shown, half off-peak |
| Sep 21 | xAI Grok 4.7 | $2 / $6 | New model at Grok 4.6’s price |
| Sep 22 | Anthropic Claude Opus 5.5 | $4 / $20 | New model below Opus 5; cache hits at 0.05× input |
| Sep 22 | OpenAI GPT-6 Sol | $2 / $10 | New model; cached input $0.20 |
| Sep 22 | OpenAI GPT-6 Luna | $0.10 / $0.50 | New model; cached input $0.01 |
| Sep 28 | Anthropic Claude Sonnet 5.5 | $2 / $10 | New model at Sonnet 5’s price |
| Sep 29 | OpenAI GPT-6.1 Sol | $2 / $10 | New model; cached input $0.10, half GPT-6 Sol’s |
02 — The launchesWhat the quarter’s new models cost
The chart ranks the quarter’s new text models by their output price. The spread runs from $0.50 to $50 per million, a hundredfold gap.
Output price per million tokens, models launched July to September 2026
Vendor pricing pages and changelogs, read October 3, 2026. Standard tier, short context. Gemini 3.7 and 3.8 Flash prices are introductory to December 31, 2026; GPT-5.6 and Gemini 3.6 Flash show launch prices; DeepSeek’s is the peak rate. Trusted-access models (GPT-5.6 Cyber, GPT-Rosalind) are omitted.03 — AnalysisFour patterns in the quarter
Successors at the same price or lower. Anthropic’s release notes put Opus 5 at Opus 4.8’s $5 / $25, then Opus 5.5 at $4 / $20 two months later. Sonnet 5.5 kept Sonnet 5’s $2 / $10, and xAI held Grok at $2 / $6 through two releases.
Cheaper cache reads. Fable 5.1 charges 0.025 times its input price for a cache read and Opus 5.5 charges 0.05 times, against the 0.1 times Anthropic charges on its other models. OpenAI’s changelog lists GPT-6.1 Sol’s cached input at $0.10, half of GPT-6 Sol’s. For agents that resend long context on every turn, these cuts can matter more than the headline rates.
Cuts soon after launch. OpenAI cut GPT-5.6 Luna by 80% and Terra by 20% on July 30, three weeks after the family launched, then put GPT-5.6 Sol on a promotional $4 / $20 on August 21, down from $5 / $30.
Time-of-day pricing. DeepSeek moved to peak and off-peak rates on August 16, with off-peak at half the peak rate; both were above the flat price they replaced. Its pricing page sets peak hours at 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. Our note on off-peak LLM pricing covers how to schedule around it.
Below its predecessor
Claude Opus 5.5 lists at $4 / $20 against Opus 5’s $5 / $25.
Fable 5.1 cache reads
Down from $1 on Fable 5, with input and output unchanged.
Gemini Flash introductory rates
3.6, 3.7 and 3.8 Flash rise from $0.75 / $3.75 to $1.50 / $7.50.
04 — The catchPrices that come with an end date
Several of the quarter’s prices are temporary. Google’s Gemini API pricing lists Gemini 3.6, 3.7 and 3.8 Flash at introductory rates through December 31, 2026, doubling on January 1, 2027. OpenAI says GPT-5.6 Sol’s promotional price is available at least through November 21, 2026. One temporary price went the other way: Anthropic made Sonnet 5’s introductory $2 / $10 permanent on August 10 and cancelled the rise it had scheduled for September 1.
A cost model built on an introductory rate breaks on a known date. Price any workload on Gemini 3.6, 3.7 or 3.8 Flash at the January rate, and treat GPT-5.6 Sol’s promotion as ending after November 21 unless OpenAI says otherwise. Our guide to budgeting on introductory pricing shows the arithmetic.
05 — Practical implicationsWhat to do with this table
For current prices across every surface, not just changes, use our maintained frontier model API price index. The August cuts and promotions tracker and the Q2 price tracker give earlier context. Teams that want model costs modelled against their real workloads can work with our AI transformation team.
06 — The re-pricingRe-pricing a workload with the table
The conclusion says to re-run last month’s token counts against the current prices. The table makes that a mechanical job if it is done in the right order, and a misleading one if it is not. Start from the usage record, not from the table: for each model you ran, pull the month’s uncached input, cached input and output token totals from the provider’s own billing export. Those three totals are the workload; the table only supplies rates.
Then find each model’s current row and, separately, the row for its successor if the quarter produced one. Opus 5 users have Opus 5.5 at $4 / $20; GPT-6 Sol users have GPT-6.1 Sol with cached input at half the price. Apply both rows to the same three totals. The difference is the saving the successor would have produced on last month’s work, before any change in quality, retries or review time, which this table does not measure and which our guide to review cost on cheaper models covers.
Finally, check every model in the plan for an end date in the table’s Change column. A row priced at an introductory rate is re-priced a second time at the rate it will carry on the day the promotion ends, where that rate is published; for GPT-5.6 Sol, assume its pre-promotion $5 / $30 unless OpenAI says otherwise. The plan keeps both numbers. A budget that carries only the introductory figure is correct until a date the table already knows.
Keep the re-pricing in the same document as the usage export, with the table’s as-of date written at the top. When the Q4 edition of this table arrives, the exercise is repeated against the same usage record, which is what turns a one-off check into a quarterly habit: the workload stays constant and only the rate column moves, so the difference between two quarters is a clean measure of what the vendors’ price changes did to your bill, as long as the model mix is held constant.
07 — The failuresFour ways a price table misleads
A dated table of list prices is a good instrument that answers one question. The illustrative mistakes below are each a case of asking it a different one.
The headline-only comparison: two models are ranked on their output price and the cache-read price is ignored, when for an agent that resends long context every turn the cache rate can decide the bill, as the patterns section shows for Fable 5.1 and Opus 5.5. The frozen promotion: a 2027 forecast is built on Gemini 3.7 Flash at $0.75 / $3.75 without the January 1 doubling the pricing page already states. The flat-rate assumption: a DeepSeek workload is costed at one price when the vendor has billed peak and off-peak rates since August 16, and the plan did not say which hours the jobs run in. The successor as a free lunch: a cheaper successor is adopted on its list price, and the review work its output needs is never measured, so the saving on the bill is real and the saving on the task is unknown.
The table answers the first three in its Change column, which records cache-read prices, end dates and DeepSeek’s peak and off-peak split, provided it is read. The fourth it cannot answer, and the method rows below say so.
A fifth, smaller misreading is worth naming: taking a list price for a price. Almost every row is a standard-tier list rate for short prompts, as the first section says; the exception is OpenAI’s Fast mode row. Volume and enterprise agreements, regional uplifts and other premium speed tiers are excluded by the method, so a team on negotiated terms should treat the table as the public reference its own rates are measured against, not as the figure on its invoice, and should write its own negotiated rate beside each row it uses.
It tells you what each vendor said it would charge, on which date, with which end date. It does not tell you what your workload costs or what the output is worth; those need your usage record and your review time.
08 — MethodMethod and as-of date
A dated log of published list-price changes for text-model APIs from five vendors, read from each vendor’s own pages.
- What was collected
- The dated changes to standard-tier text-model API prices between July 1 and September 30, 2026 that we found in the five vendors’ changelogs and pricing pages, including new models at their launch price: 28 rows.
- Sources
- Anthropic’s release notes and pricing page; OpenAI’s API changelog and pricing page; the Gemini API changelog and pricing page; DeepSeek’s update log and pricing page; xAI’s release notes. No routing marketplaces or third-party price lists.
- As-of date
- October 3, 2026.
- Exclusions
- Image, video, voice and transcription models; enterprise and volume pricing; regional and data-residency uplifts; premium speed tiers other than OpenAI’s July 30 Fast mode, which replaced an existing tier; vendors outside these five.
- Limitations
- The old prices behind OpenAI’s July 30 cuts, DeepSeek’s August 16 and September 10 changes and Gemini 3.6 Flash’s launch are no longer on the vendors’ current pages; they were read from Internet Archive copies of those pages.
- Refresh
- A Q4 edition follows the quarter. Errors found in this table are corrected in place with a dated note.
Re-price your workloads on the new rates
If you set model budgets in June, at least one of your rows has probably moved. Re-run last month’s token counts against the current prices, check whether a cheaper successor now exists for each model you use, and note every price that expires.