AI DevelopmentPricing Tracker7 min readPublished October 3, 2026

Twenty-eight dated price events, July to September 2026

AI Model API Price Changes, Q3 2026: Five Labs, One Table

API price changes from Anthropic, OpenAI, Google, DeepSeek and xAI, July to September 2026: launches, cuts, a rise and promotions per million tokens, dated.

DA
Digital Applied Team
Research and practical guidance
CoverageOctober 3, 2026

Between July 1 and September 30, 2026, Anthropic, OpenAI, Google, DeepSeek and xAI made at least 28 changes to the prices of their text-model APIs, counting new models and their launch prices. Most new flagships arrived at or below the price of the model they replaced, one vendor raised its prices, and several prices now carry an end date. Every row below is dated from the vendor’s own changelog or pricing page.

Key takeaways
  1. 01
    Launches held the lineOpus 5, Sonnet 5.5 and Grok 4.6 and 4.7 matched their predecessors. Opus 5.5 came in 20% below Opus 5.
  2. 02
    Caching got cheaperFable 5.1, Opus 5.5 and GPT-6.1 Sol all cut the price of reading cached input.
  3. 03
    Prices with expiry datesGemini 3.6, 3.7 and 3.8 Flash double on January 1; GPT-5.6 Sol’s promotion runs at least to November 21.
  4. 04
    Time of day now countsSince August 16, DeepSeek bills peak and off-peak rates, with off-peak at half the peak rate. Both sit above its old flat price.

01 — The dataThe quarter’s price changes, dated

Prices are US dollars per million input and output tokens at the standard tier, for prompts under any long-context threshold. Where a vendor’s current page no longer shows an old price, it was read from an Internet Archive copy of that vendor’s own pricing page.

Sources: Anthropic, OpenAI, Google Gemini API, DeepSeek and xAI changelogs and pricing pages, read October 3, 2026.
DateModelPriceChange
Jul 8xAI Grok 4.5$2 / $6New model
Jul 9OpenAI GPT-5.6 Sol$5 / $30New model at GPT-5.5’s price
Jul 9OpenAI GPT-5.6 Terra$2.50 / $15New model
Jul 9OpenAI GPT-5.6 Luna$1 / $6New model
Jul 21Google Gemini 3.6 Flash$1.50 / $7.50Generally available; same input price as 3.5 Flash ($1.50 / $9), lower output
Jul 21Google Gemini 3.5 Flash-Lite$0.30 / $2.50Generally available
Jul 24Anthropic Claude Opus 5$5 / $25New model at the same price as Opus 4.8
Jul 30OpenAI GPT-5.6 Luna$0.20 / $1.20Price cut by 80%, from $1 / $6
Jul 30OpenAI GPT-5.6 Terra$2 / $12Price cut by 20%, from $2.50 / $15
Jul 30OpenAI Fast mode2× standardReplaces Priority Processing; for GPT-5.6 Sol it costs twice standard
Aug 7OpenAI GPT-5.6 Cyber$12.50 / $75New model, approved Daybreak program access only
Aug 10Anthropic Claude Sonnet 5$2 / $10Introductory price made standard; the Sep 1 rise to $3 / $15 cancelled
Aug 12xAI Grok 4.6$2 / $6New model at Grok 4.5’s price
Aug 13Google Gemini 3.7 Flash$0.75 / $3.75New model at an introductory price to Dec 31; $1.50 / $7.50 from Jan 1, 2027
Aug 13Google Gemini 3.6 Flash$0.75 / $3.75Cut to an introductory price to Dec 31; back to $1.50 / $7.50 from Jan 1, 2027
Aug 16DeepSeek API$0.44 / $1.32 peakPeak and off-peak billing replaces a flat rate; V4 Flash rises from $0.14 / $0.28 to $0.44 / $1.32 at peak, off-peak is half
Aug 21OpenAI GPT-5.6 Sol$4 / $20Promotional price, input 20% and output 33% lower, at least to Nov 21
Sep 1Anthropic Claude Fable 5.1$10 / $50New model; cache reads $0.25, down from $1 on Fable 5
Sep 2Google Gemini 3.8 Flash$0.75 / $3.75New model at an introductory price to Dec 31; $1.50 / $7.50 from Jan 1, 2027
Sep 3OpenAI GPT-6 Astra$10 / $50New model; $20 / $75 for prompts above 272K tokens
Sep 8OpenAI GPT-Rosalind$5 / $25New model, trusted access only; billing starts Oct 5
Sep 10DeepSeek V4.1 Flash$0.30 / $1.20Replaces V4 Flash ($0.44 / $1.32 peak); peak rates shown, half off-peak
Sep 21xAI Grok 4.7$2 / $6New model at Grok 4.6’s price
Sep 22Anthropic Claude Opus 5.5$4 / $20New model below Opus 5; cache hits at 0.05× input
Sep 22OpenAI GPT-6 Sol$2 / $10New model; cached input $0.20
Sep 22OpenAI GPT-6 Luna$0.10 / $0.50New model; cached input $0.01
Sep 28Anthropic Claude Sonnet 5.5$2 / $10New model at Sonnet 5’s price
Sep 29OpenAI GPT-6.1 Sol$2 / $10New model; cached input $0.10, half GPT-6 Sol’s

02 — The launchesWhat the quarter’s new models cost

The chart ranks the quarter’s new text models by their output price. The spread runs from $0.50 to $50 per million, a hundredfold gap.

Output price per million tokens, models launched July to September 2026

Vendor pricing pages and changelogs, read October 3, 2026. Standard tier, short context. Gemini 3.7 and 3.8 Flash prices are introductory to December 31, 2026; GPT-5.6 and Gemini 3.6 Flash show launch prices; DeepSeek’s is the peak rate. Trusted-access models (GPT-5.6 Cyber, GPT-Rosalind) are omitted.
Claude Fable 5.1
$50
GPT-6 Astra
$50
GPT-5.6 Sol
$30
Claude Opus 5
$25
Claude Opus 5.5
$20
GPT-5.6 Terra
$15
Claude Sonnet 5.5
$10
GPT-6 Sol
$10
GPT-6.1 Sol
$10
Gemini 3.6 Flashat launch
$7.50
Grok 4.5, 4.6 and 4.7
$6
GPT-5.6 Lunaat launch
$6
Gemini 3.7 and 3.8 Flashintroductory
$3.75
Gemini 3.5 Flash-Lite
$2.50
DeepSeek V4.1 Flashpeak
$1.20
GPT-6 Luna
$0.50

03 — AnalysisFour patterns in the quarter

Successors at the same price or lower. Anthropic’s release notes put Opus 5 at Opus 4.8’s $5 / $25, then Opus 5.5 at $4 / $20 two months later. Sonnet 5.5 kept Sonnet 5’s $2 / $10, and xAI held Grok at $2 / $6 through two releases.

Cheaper cache reads. Fable 5.1 charges 0.025 times its input price for a cache read and Opus 5.5 charges 0.05 times, against the 0.1 times Anthropic charges on its other models. OpenAI’s changelog lists GPT-6.1 Sol’s cached input at $0.10, half of GPT-6 Sol’s. For agents that resend long context on every turn, these cuts can matter more than the headline rates.

Cuts soon after launch. OpenAI cut GPT-5.6 Luna by 80% and Terra by 20% on July 30, three weeks after the family launched, then put GPT-5.6 Sol on a promotional $4 / $20 on August 21, down from $5 / $30.

Time-of-day pricing. DeepSeek moved to peak and off-peak rates on August 16, with off-peak at half the peak rate; both were above the flat price they replaced. Its pricing page sets peak hours at 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. Our note on off-peak LLM pricing covers how to schedule around it.

Cheaper
Below its predecessor
−20%Opus 5.5

Claude Opus 5.5 lists at $4 / $20 against Opus 5’s $5 / $25.

Sep 22
Cache
Fable 5.1 cache reads
$0.25per million

Down from $1 on Fable 5, with input and output unchanged.

Sep 1
Expiry
Gemini Flash introductory rates
2×on Jan 1

3.6, 3.7 and 3.8 Flash rise from $0.75 / $3.75 to $1.50 / $7.50.

Dec 31 end

04 — The catchPrices that come with an end date

Several of the quarter’s prices are temporary. Google’s Gemini API pricing lists Gemini 3.6, 3.7 and 3.8 Flash at introductory rates through December 31, 2026, doubling on January 1, 2027. OpenAI says GPT-5.6 Sol’s promotional price is available at least through November 21, 2026. One temporary price went the other way: Anthropic made Sonnet 5’s introductory $2 / $10 permanent on August 10 and cancelled the rise it had scheduled for September 1.

Budget on the later price

A cost model built on an introductory rate breaks on a known date. Price any workload on Gemini 3.6, 3.7 or 3.8 Flash at the January rate, and treat GPT-5.6 Sol’s promotion as ending after November 21 unless OpenAI says otherwise. Our guide to budgeting on introductory pricing shows the arithmetic.

05 — Practical implicationsWhat to do with this table

Running Opus 5 on long agent tasks
Test Opus 5.5: lower list price and cheaper cache hits
Anthropic
Heavy repeated context on any model
Compare cache-read prices, not just input and output
All vendors
Budgeting 2027 on Gemini Flash
Use the January 1 rates in the plan
Google
Batch jobs that can wait
Schedule off-peak on DeepSeek, or use batch tiers
DeepSeek and others

For current prices across every surface, not just changes, use our maintained frontier model API price index. The August cuts and promotions tracker and the Q2 price tracker give earlier context. Teams that want model costs modelled against their real workloads can work with our AI transformation team.

06 — The re-pricingRe-pricing a workload with the table

The conclusion says to re-run last month’s token counts against the current prices. The table makes that a mechanical job if it is done in the right order, and a misleading one if it is not. Start from the usage record, not from the table: for each model you ran, pull the month’s uncached input, cached input and output token totals from the provider’s own billing export. Those three totals are the workload; the table only supplies rates.

Then find each model’s current row and, separately, the row for its successor if the quarter produced one. Opus 5 users have Opus 5.5 at $4 / $20; GPT-6 Sol users have GPT-6.1 Sol with cached input at half the price. Apply both rows to the same three totals. The difference is the saving the successor would have produced on last month’s work, before any change in quality, retries or review time, which this table does not measure and which our guide to review cost on cheaper models covers.

Finally, check every model in the plan for an end date in the table’s Change column. A row priced at an introductory rate is re-priced a second time at the rate it will carry on the day the promotion ends, where that rate is published; for GPT-5.6 Sol, assume its pre-promotion $5 / $30 unless OpenAI says otherwise. The plan keeps both numbers. A budget that carries only the introductory figure is correct until a date the table already knows.

Keep the re-pricing in the same document as the usage export, with the table’s as-of date written at the top. When the Q4 edition of this table arrives, the exercise is repeated against the same usage record, which is what turns a one-off check into a quarterly habit: the workload stays constant and only the rate column moves, so the difference between two quarters is a clean measure of what the vendors’ price changes did to your bill, as long as the model mix is held constant.

07 — The failuresFour ways a price table misleads

A dated table of list prices is a good instrument that answers one question. The illustrative mistakes below are each a case of asking it a different one.

The headline-only comparison: two models are ranked on their output price and the cache-read price is ignored, when for an agent that resends long context every turn the cache rate can decide the bill, as the patterns section shows for Fable 5.1 and Opus 5.5. The frozen promotion: a 2027 forecast is built on Gemini 3.7 Flash at $0.75 / $3.75 without the January 1 doubling the pricing page already states. The flat-rate assumption: a DeepSeek workload is costed at one price when the vendor has billed peak and off-peak rates since August 16, and the plan did not say which hours the jobs run in. The successor as a free lunch: a cheaper successor is adopted on its list price, and the review work its output needs is never measured, so the saving on the bill is real and the saving on the task is unknown.

The table answers the first three in its Change column, which records cache-read prices, end dates and DeepSeek’s peak and off-peak split, provided it is read. The fourth it cannot answer, and the method rows below say so.

A fifth, smaller misreading is worth naming: taking a list price for a price. Almost every row is a standard-tier list rate for short prompts, as the first section says; the exception is OpenAI’s Fast mode row. Volume and enterprise agreements, regional uplifts and other premium speed tiers are excluded by the method, so a team on negotiated terms should treat the table as the public reference its own rates are measured against, not as the figure on its invoice, and should write its own negotiated rate beside each row it uses.

What the table is for, in one sentence

It tells you what each vendor said it would charge, on which date, with which end date. It does not tell you what your workload costs or what the output is worth; those need your usage record and your review time.

08 — MethodMethod and as-of date

Methodology

A dated log of published list-price changes for text-model APIs from five vendors, read from each vendor’s own pages.

What was collected
The dated changes to standard-tier text-model API prices between July 1 and September 30, 2026 that we found in the five vendors’ changelogs and pricing pages, including new models at their launch price: 28 rows.
Sources
Anthropic’s release notes and pricing page; OpenAI’s API changelog and pricing page; the Gemini API changelog and pricing page; DeepSeek’s update log and pricing page; xAI’s release notes. No routing marketplaces or third-party price lists.
As-of date
October 3, 2026.
Exclusions
Image, video, voice and transcription models; enterprise and volume pricing; regional and data-residency uplifts; premium speed tiers other than OpenAI’s July 30 Fast mode, which replaced an existing tier; vendors outside these five.
Limitations
The old prices behind OpenAI’s July 30 cuts, DeepSeek’s August 16 and September 10 changes and Gemini 3.6 Flash’s launch are no longer on the vendors’ current pages; they were read from Internet Archive copies of those pages.
Refresh
A Q4 edition follows the quarter. Errors found in this table are corrected in place with a dated note.
Next step

Re-price your workloads on the new rates

If you set model budgets in June, at least one of your rows has probably moved. Re-run last month’s token counts against the current prices, check whether a cheaper successor now exists for each model you use, and note every price that expires.

AI cost engineering

Know what your models will cost next quarter

Digital Applied models AI spend against your real token usage, finds the cheaper route for each workload and flags prices that are about to change.

Usage-based modelsRoute comparisonsExpiry tracking
Before you budget

Check four prices

  • →Input and output
  • →Cache reads
  • →Long-context rates
  • →Expiry dates
Questions and answers

Practical questions

$4 per million input tokens and $20 per million output tokens, with cache hits at 0.05 times the input price, per Anthropic’s pricing page as of October 3, 2026. Claude Opus 5 costs $5 / $25.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source