Background
Changelog

What changed,
and what it cost.

Every product update and price movement — explained in words, not API docs. Because a 400% price increase means more when you see it as a novel going from $0.35 to $1.33.

This page is a dated record. Prices quoted inside an entry are the prices as recorded on that entry’s date, not necessarily today’s rates. Current prices live on the comparison table, checked nightly.

DeepSeek gets a conservative peak-rate rule, and Groq returns to its current production lineup

DeepSeek now publishes two standard on-demand windows rather than one flat price. TokenScale uses the highest published rate as the comparison figure and labels it as peak: V4 Flash is $0.44 input and $1.32 output per million tokens; V4 Pro is $1.32/$3.96. The exact off-peak rates are 50% lower outside 01:00–04:00 and 06:00–10:00 UTC. We do not invent an average because each customer's traffic mix is different.

Groq's retired Llama 3.1 8B and Llama 3.3 70B developer rows remain in dated history, but no longer occupy today's board. Small now tracks production GPT-OSS 20B at $0.075/$0.30. Mid and Flagship both track production GPT-OSS 120B at $0.15/$0.60 because Groq currently has only two public production text models. TokenScale does not fill the third tier with a Preview model.

Two official price notes reached the board without becoming invented history

Alibaba still marks Qwen3.7-Max as a limited-time 50% promotion, but its public pricing page does not supply a usable start and end date. TokenScale therefore continues to record the published $2.50 input and $7.50 output list rate, with no backfilled promotion window.

Sakana also added an untracked Namazu model at $0.95 input and $4 output per million tokens. Fugu Ultra remains the tracked flagship at $5/$30. Both facts now sit on the public watchlist for a dated follow-up; neither was turned into an automatic model swap or a made-up price point.

Our own reliability page understated its gaps. Here is the correction, and the guard that stops it happening again.

The reliability page exists to count the nights the record missed. From 4 to 9 August it counted them wrong, and wrong in its own favour: it showed 14 carried-forward nights while the price record held 17, and it left the August run marked "ongoing" five days after that run had ended. The nights of 4, 6 and 7 August were never disclosed at all.

What actually happened, from the data. The run that began on 1 August went one night longer than published, through 4 August, and ended on 5 August, when a pass verified 39 of the 66 tracked tiers and held the other 27. A second run followed on 6 and 7 August with no verification pass at all. The 8 August run reached 3 tiers before it stopped. The board has not had a complete nightly pass since 31 July, so that run is left open on the reliability page rather than quietly closed.

The guard matters more than the fix. Counting by hand is what failed, so those counters are no longer typed: they are computed from the price record itself on every build. Each row in the gap table now declares which nights it covers, and the build fails if a night in the record has gone undisclosed for more than three days. The page can still trail the record by up to three nights, and while it does it says so on its face instead of rounding down. The corrected record: 82 nights since 19 May, 63 fully verified, 2 partly verified, 17 carried forward.

We audited the site against its own record, and found four things. All fixed, all published.

First: the live version file, BUILD.txt, could serve a cached copy up to five days stale (it reported v514 while v518 was live). It now ships with a no-store header, so neither browsers nor the CDN may cache it. Second: the nights of 1, 2 and 3 August were never verified. The nightly kept appending data, but no verification pass ran and no provider was stamped after 31 July. Those three nights now stand in the record as carried forward, permanently.

Third, the big one: the methodology page said history is append-only, then said grey carried-forward days "turn solid" once the next check lands. Those can't both be true, and the charts really did repaint held days as verified. The append-only rule wins. Carried-forward nights are now flagged in the data itself, back to May, and the flag never clears: a grey day stays grey, because a later check cannot confirm a night it didn't see. Fourth: the Basic page's no-JavaScript table and its "Rates verified" caption were frozen at 1 July; both are regenerated and stamped 31 July, the true date of the last full verification.

Two additions come with this. A new reliability page lists every night the record missed since 19 May, with the reason: fourteen of seventy-seven, counted honestly. And from 5 August the site's pages are submitted to the Internet Archive nightly by a watchdog running outside the machine the nightly runs on, so there is a third-party, timestamped copy of what this site said that we don't control and can't edit. The full story is in the journal.

The nightly stalled for three days — so we hardened it, and made every carried-forward day visible

Between 13 and 17 July the tracker stopped writing new data, and it took three nights to notice. The cause was mundane: a macOS update reset a permission the background job needs to read its own files, and it failed — quietly — five times in a row. That silence is what we fixed first. The job now raises an alarm the moment it fails, names the actual cause instead of shrugging, and runs on a footing an OS update can’t knock out again. A three-day blackout can’t happen unseen twice.

TokenScale is built to verify every price against the provider’s official page every night. On any night that check can’t complete, the price is carried forward from the last verification so the chart stays whole — and here’s the part worth keeping: those carried-forward days are now drawn in grey, on every chart and table, with each provider stamped at the date it was last actually checked. Every number is real and traces to an official page; nothing was invented, and nothing was lost in the outage. A held price simply shouldn’t wear the confidence of a freshly-checked one — so now it doesn’t.

The methodology page spells this out in full. It’s the same standard we hold providers to: don’t hide the number, and don’t dress an estimate up as a measurement. When we can’t verify a night, we’d rather show you the grey than pretend it’s black and white.

Fable 5 returns — and this time it keeps the flagship slot

Some models get a launch. Fable 5 got a saga. Anthropic's Mythos-class flagship held our Pro slot for exactly 48 hours in June — 10th to 12th — before a US export-control directive suspended it worldwide, and the board reverted to Opus 4.8. We left those two $10/$50 points in the history because they happened. Yesterday's audit confirmed it: Fable 5 is back on Anthropic's price list at $10/$50, and as of today it retakes the Anthropic Pro slot. Opus 4.8 gets the Sonnet 4.6 treatment — it moves to the all-models view at $5/$25, still listed, still available, nothing deleted.

Read the chart honestly: the Pro line shows the June blip, three weeks of Opus at $5/$25, and a step to $10/$50 today. That step is a model swap, badged as one — the slot changed which model it tracks; nobody's Opus bill doubled. In real terms, the priciest way to write The Hobbit on this board is now about $6.36 (127K output tokens at $50/M), against $0.08 on the cheapest tier. And yes, we note the poetry: the model that now tops the board is the same one that spent yesterday auditing it.

The July drift audit — we re-checked all 66 tiers, and here's everything that was wrong

Once in a while the right move is to stop trusting your own machinery and check everything by hand. On 1 July we ran a full drift audit: every one of the 22 providers × 3 tiers on this board, re-verified against the provider's own pricing page, in one sweep, by Claude Fable 5. The score: 40 of 66 tiers checked out exactly. The other 26 are why this entry exists — and in the spirit of this site, we're publishing the misses, not just the fixes.

The worst find is a billing trap, not a typo. Our xAI lite tier still showed Grok 4.1 Fast at $0.20/$0.50. That model was retired on 15 May — but the old API slugs don't error. They silently redirect to Grok 4.3 and bill at $1.25/$2.50, roughly 6× the price you thought you were paying. If you had an agent pointed at that slug, your invoice already knows. The slot now shows Grok Build 0.1 ($1/$2) — xAI no longer sells a sub-dollar text model at all.

Two models on our board never existed. "Kimi K2.6 Turbo" and "MiniMax M2.7 Pro" appear in no official price list — they were plausible-sounding names that slipped in during fast-moving update nights and survived because their prices looked reasonable. A third, Llama 4 Behemoth, was announced in 2025 and never publicly shipped — we were quoting real prices for a ghost. All three are gone: the Kimi slot now tracks K2.7 Code HighSpeed ($1.90/$8), MiniMax gets its actual new flagship M3 ($0.30/$1.20 — flagship capability at lite money, the cheapest flagship tier on the board), and Meta's third slot falls back to Maverick, the biggest Llama you can actually buy.

Then there's the quiet Llama exodus. Eight of our 66 tiers pointed at Llama models their hosts no longer serve: Fireworks has dropped every serverless Llama, Cerebras retired both of its (one the day after our last hand-check), SambaNova's 405B has been gone for over a year, Together delisted the 3B and 405B, DeepInfra dropped the 405B, and Groq's DeepSeek R1 distill was shut down last October. Groq's two remaining Llamas die 16 August — diarised. The GPT-OSS models have quietly become the standard budget tier across the fast-inference hosts, and the board now reflects that. Mistral, meanwhile, had moved a full generation (Small 4, Medium 3.5, Large 3) — and yes, Mistral Large 3 now costs less than Mistral Medium 3.5. Their own FAQ still contradicts their own API table; we've flagged it and gone with the table. Alibaba's whole tracked lineup had been reclassified "legacy, not recommended" by Alibaba itself — replaced with the 3.6/3.7 generation. And Zhipu's "GLM-5-Air" was our mislabel: the $0.20/$1.10 model is GLM-4.5-Air. Right price, wrong name, ours.

The honest bit about our own machinery. The nightly price check did its job — it verified prices. What it couldn't see is that some of those prices belonged to models that had been retired, delisted, or never existed: it was faithfully price-checking ghosts. Three providers' "verified" stamps had quietly frozen at 26 May while the footer said "verified nightly." That's drift, it was ours, and it's the exact failure mode this site exists to catch in others. Fable found it, Fable fixed it, and model-existence checks are joining the nightly routine so the graveyard tends itself. Speaking of which: today's departures — the dead Llamas, the phantom Turbo and Pro, the never-born Behemoth — are all being laid to rest with dates and final prices in the Model Graveyard →

What survived the audit untouched: Google, OpenAI, Anthropic, DeepSeek, Cohere, Perplexity, OpenRouter and Sakana — every tier exact. AWS Bedrock remains the one provider we could not re-verify from source (their pricing page won't render without JavaScript); it's consistent with Anthropic's own global-endpoint pricing, and it keeps its 23 June stamp until we can read it from AWS itself. That's the whole ledger. Next full audit: when the data earns one — the nightly watch continues in the meantime, every night, at Price Moves →

Claude Sonnet 5 takes the mid tier — same sticker price, a heavier tokenizer

Anthropic launched Claude Sonnet 5 on 30 June. It steps into TokenScale's Anthropic Mid slot in place of Sonnet 4.6, at the same standard list price — $3/M in, $15/M out — and Anthropic says it closes much of the gap to Opus. Sonnet 4.6 isn't deleted: it moves to the "all models" view, still listed and still priced, the way we kept Opus alongside Fable rather than rewriting the record.

Two things worth knowing before you point an agent at it. First, a launch discount: through 31 August, Sonnet 5 runs at $2/M in, $10/M out, then reverts to $3/$15. We track the standing list price as canonical, so the headline number stays $3/$15 and the discount is a temporary window, not history. Second — and very much our beat — Sonnet 5 ships a new tokenizer that turns the same text into roughly 30% more tokens. The per-token price is unchanged, but the same email or novel can cost about a third more to run. That's exactly the kind of hidden shift a sticker price hides and a content-size lens makes visible.

GPT-5.6 lands in preview — Sol, Terra and Luna, gated to a handful of partners

OpenAI unveiled the GPT-5.6 family on 26 June — three models named Sol, Terra and Luna, a new naming scheme that replaces the old mini/nano tiers with capability tiers inside one generation. Access is the story in itself: it's a limited preview reachable only through an OpenAI account rep, after OpenAI shared the models and its release plan with the US government first. General availability is promised "in the coming weeks."

Verified list pricing, per million tokens: Luna $1/$6, Terra $2.50/$15, Sol $5/$30 — Sol matches GPT-5.5's flagship rate, and Terra lands exactly on today's GPT-5.4 mid-tier. Because you can't yet buy these off the shelf, TokenScale keeps GPT-5.5 and GPT-5.4 as the live Lite/Mid/Pro tiers and lists Sol, Terra and Luna in OpenAI's "all models" view, marked preview. We'll promote them to the headline slots the moment they're generally available — the same way we handled Claude Fable 5's brief appearance rather than rewriting the board around a model most people can't reach.

We turned on anonymous engagement measurement — here's exactly what it does

TokenScale now records a small set of anonymous, cookieless engagement events — things like "the page was scrolled past the pricing", "the calculator was used", or "a provider was viewed" — so we can see which parts actually help and build the right things next. In the spirit of the site, we're spelling out precisely what this is and isn't.

What we don't do: no cookies, no IP address stored, no fingerprinting, and the text you paste into the calculator is never sent — it stays in your browser, exactly as before. To follow a single visit's flow we tag its events with a random ID that lives only in your browser's memory, is recreated on every page load, and vanishes when you close the tab; it can't be linked across visits or back to you. We also honour "Do Not Track" and Global Privacy Control — switch either on and we collect nothing at all.

The full detail is on the privacy page →

GLM-5.2 takes a flagship tier for open-weight money

Zhipu launched GLM-5.2 on 16 June — the first MIT-licensed, 1M-context model to hold a flagship tier. TokenScale's Zhipu GLM flagship slot moves to $1.40/M in, $4.40/M out, against GPT-5.5's $5/$30 at the same context. The +43% over GLM-5.1 is a generational step up, not a repricing.

+43%Flagship in & out · $0.98→$1.40 / $3.08→$4.40
$1.40/Minput · vs GPT-5.5 $5/M
MITlicensed · 1M context

A five-tier reshuffle in one night

13 June re-priced five tiers at once, in both directions. Cerebras jumped +299% while AWS Bedrock cut hosted Llama 3.1 405B output from $16 to $2.40 — the night's biggest drop, the opposite direction on the same run. Mistral, Qwen and an OpenRouter route moved too. A clean illustration of how fast the floor shifts.

5tiers re-priced in one night
+299%Cerebras
−85%AWS Llama 3.1 405B out · $16→$2.40

Correction: Anthropic's Pro tier is back to Opus 4.8 ($5/$25)

The 10 June note below recorded Anthropic's Pro slot moving to Fable 5 at $10/$50. That didn't hold. Between 10–12 June the Pro tier round-tripped — $5→$10→$5 in, $25→$50→$25 out — settling back at Opus 4.8 ($5/$25). TokenScale tracks Opus 4.8 as Anthropic's flagship Pro tier; Fable 5 lives in the all-models view, not the headline slot. We're leaving the 10 June entry up as a record of the round-trip rather than rewriting it.

round-trip$5→$10→$5 in · $25→$50→$25 out
$5/$25Opus 4.8 · Pro tier now

DeepSeek V4 Pro: another −75% cut

DeepSeek cut its Mid/Pro tier (V4 Pro) by 75% on 11 June — $1.74→$0.435 in, $3.48→$0.87 out. The relentless undercutting that's kept the Novel Index floor near half a cent.

−75%Mid/Pro · V4 Pro
$0.435/Minput · was $1.74
$0.87/Moutput · was $3.48

Claude Fable 5 lands: Anthropic's Pro slot just doubled

Anthropic launched Claude Fable 5 on 9 June, its first generally available Mythos-class model: a tier above Opus. TokenScale's Anthropic Pro slot now maps to Fable 5 (high) at $10/M in, $50/M out. That makes it the most expensive model on the board, at exactly 2× Opus 4.8. Reading The Hobbit now costs $1.27 on Anthropic's top tier, up from $0.63.

+100%Pro tier in & out · $5→$10 / $25→$50
$1.27The Hobbit, input · was $0.63
$50/Moutput · priciest rate tracked

Opus 4.8 isn't gone: it stays in the all-models view alongside Haiku, Sonnet and Fable. Batch pricing halves Fable's rates to $5/$25, and cached reads drop input to $1/M. Worth knowing before you point an agent at a novel.

Five new labs — TokenScale now tracks 21 providers

DeepSeek used to be the only Chinese lab on the board. That stopped making sense. Added four open-weight frontier labs — Alibaba Qwen, Zhipu GLM, Moonshot Kimi and MiniMax — plus Meta's own first-party Llama API, so you no longer have to price Llama through a reseller.

21providers tracked · was 16
63price points verified nightly
5new labs added

Four of the five are open-weight and price like it — often an order of magnitude below a US frontier flagship. A handful of top-tier prices that aren't public yet are careful estimates, marked for correction. The full story is in the journal →

Price Moves — every change, now shareable

The month's biggest swings now live on their own page — and each move exports as a branded card you can drop straight into a thread. A percentage tells a developer something; a whole novel going from $1.91 to $0.32 tells everyone.

−83%Grok mid · biggest cut
+275%OpenAI Pro · biggest rise
a whole novel on DeepSeek V4

Each card is drawn on your own device — no server, no tracking, pure static HTML and a little canvas. See all seven moves →

Gemini Flash 3.5 quietly got 5× more expensive

Google updated pricing on their mid-tier Gemini model. Silver (Flash-Lite) stayed the same. Gold (Flash 3.5) did not.

+400%Flash 3.5 input
+260%Flash 3.5 output
Flash-Lite (unchanged)
Rate changes per million tokens
Model Input before Input after Change
Flash-Lite 2.5 · Silver $0.10 $0.10
Flash 3.5 · Gold $0.30 $1.50 +400% ↑
🧙 The Hobbit — 95,356 words — what it now costs to process
Model Before After Difference
Flash-Lite 2.5 · Silver $0.06 $0.06
Flash 3.5 · Gold $0.35 $1.33 +$0.98 · +380% ↑
Gap between tiers $0.29 $1.27 4.4× wider
The irony
TokenScale launched saying "The Hobbit = $0.06 on Gemini Flash-Lite." That's still true. But the next model up — Flash 3.5 — had its Hobbit cost jump from $0.35 to $1.33 almost immediately after launch. The gap between Silver and Gold went from $0.29 to $1.27 overnight. If you were building on Flash 3.5, you just got a 380% bill increase with no warning. This is exactly what TokenScale exists to catch.
What this means practically
If you're sending novel-length context (100K tokens) to Gemini, Flash-Lite is still the smart choice — unchanged, and still the cheapest full-context model in the comparison. Flash 3.5 now costs 22× more for the same input. The tier gap matters more than ever.

We caught a pricing error on our own launch morning

Hours before posting to Hacker News, we found that our hero number was quoting input cost only. Here's what we fixed — and why it made for a better story.

The correction
What we said What it actually was Corrected to
The Hobbit on Gemini Flash $0.04 (input only) $0.06 (total)
Input cost $0.01 $0.01
Output cost missing $0.05
Correct total $0.04 ✗ $0.06 ✓
The lesson
The very thing TokenScale is built to prevent — quoting only input cost and missing output — was in our own marketing copy. Caught and corrected before the HN post went live. The distinction between input and output pricing is more important than it looks. Output tokens are usually 3–5× more expensive per token, and most real conversations generate far more output than people expect.

Quiz button crash — caught one day before launch

The "Which model should I use?" quiz was silently failing on first run. A missing DOM element meant the result screen crashed before anyone could see a recommendation.

Why this matters
The quiz is the main personalisation hook — it routes new visitors to the right provider before they see pricing. A silent crash on the most common recommendation (Anthropic) would have cost us conversions on HN launch day without us ever knowing.

TokenScale ships — 16 providers, one page

The first public version. A single HTML file, no backend, no sign-up. Pricing for 16 AI providers expressed in content you recognise.

Pricing spread at launch — mid-tier input cost per million tokens
Provider Model $/M input Hobbit cost
DeepSeekV4 Flash$0.14$0.02
GroqLlama 3.1 8B$0.05$0.01
GeminiFlash-Lite 2.5$0.10$0.06
MistralSmall$0.10$0.01
AnthropicClaude Haiku 4.5$1.00$0.13
OpenAIGPT-5.4$2.50$0.32
OpenAIGPT-5.5$5.00$0.63
Spread cheapest → most expensive35× apart
The insight that launched this
"$5 per million tokens" tells you nothing. "The Hobbit costs $0.06 on Gemini Flash-Lite, and $0.63 on GPT-5.5" tells you everything. TokenScale was built to make that translation automatic — for any content size, across all 22 providers, verified nightly.