Anthropic still sells four tiers at once — Haiku 4.5, Sonnet 5, Opus 5 and Fable 5 — spanning 10× on input price and 50× end-to-end. Treating that menu as a single choice is how teams overpay. Treating it as a scheduling problem is how platforms win. Verified prices, honest positioning, a decision flow, and a live cost simulator below.
Updated for Claude Opus 5 (launched Jul 24, 2026): same $5/$25 sticker as Opus 4.8, a claimed >2× jump on Frontier-Bench, thinking on by default, an effort ladder up to max, and the freshest knowledge cutoff in the lineup. Every spec, chart and simulator column below has been re-verified against Anthropic's docs on Aug 4, 2026 — flip “show what changed” in the spec sheet to see the diff against the July guide.
Prices from Anthropic’s official pricing page, retrieved Aug 4, 2026. Sonnet 5 introductory pricing runs through August 31, 2026. The rhyme survives the Opus refresh: Opus 5 input still costs exactly what Haiku 4.5 output costs — Anthropic changed the model, not the ladder. The spacing is deliberate; the capability behind rung three is not what it was in July.
Anthropic doesn’t ship a single flagship. It ships a price-segmented ladder, and every tier has a legitimate job. What changed on July 24: the third rung now carries near-frontier intelligence, which quietly rewrites the job descriptions above and below it.
The fastest Claude, with near-frontier intelligence. The only current model on the previous tokenizer and a 200k window — and priced for anything you run millions of times a day.
The best balance of speed and intelligence in the lineup — the most agentic Sonnet yet, built to plan and drive tools autonomously. Launched June 30, 2026.
A step change, not an increment: close to Fable 5’s frontier intelligence at half the price, and the new state of the art on Frontier-Bench and GDPval-AA. Thinking is on by default, the effort ladder runs low → max, and it self-verifies without being told. New default on Claude Max. Opus 4.8 · May 28 · correctness tier
Next-generation intelligence for long-running agents — the generally available half of the Mythos-class tier, shipped with dual-use safety classifiers (Mythos 5 is the invitation-only variant, via Project Glasswing). Adaptive thinking is always on. GA June 9, 2026.
Everything below is from Anthropic’s official model and pricing documentation, re-verified Aug 4, 2026. The rows most people skip — cache-read rates and the tokenizer generation — are still the ones that decide real-world cost. New this revision: flip the toggle to see exactly what the Opus 5 launch changed.
| Spec | Haiku 4.5 | Sonnet 5 | Opus 5 | Fable 5 |
|---|---|---|---|---|
| Positioning | Fastest, near-frontier intelligence | Best speed-to-intelligence balance | Complex agentic coding & enterprise worknear-Fable intelligence at half the price | Next-gen intelligence for long-running agents |
| API model ID | claude-haiku-4-5 | claude-sonnet-5 | claude-opus-5 claude-opus-4-8 | claude-fable-5 |
| Input / output | $1 / $5 | $2 / $10intro to Aug 31 · then $3 / $15 | $5 / $25 UNCHANGEDfast mode $10 / $50 · API-only preview, ~2.5× speed | $10 / $50 |
| Cache read (0.1×) | $0.10 | $0.20then $0.30 | $0.50 UNCHANGEDmin cacheable prompt now 512 tok 1,024 | $1.00 |
| Batch (50% off) | $0.50 / $2.50 | $1 / $5then $1.50 / $7.50 | $2.50 / $12.50 | $5 / $25 |
| Context window | 200k | 1Mstandard pricing across full window | 1Mdefault and maximum — no smaller variant | 1Mstandard pricing across full window |
| Max output | 64k | 128k300k via Batch API beta | 128k300k via Batch API beta | 128k |
| Thinking | Extended | Adaptive | Adaptive · on by default off unless enabledmodel decides when & how much to think, per turn | Adaptive · always on |
| Effort | — | Highdefault on API & Claude Code | low · medium · high · xhigh · max high on all surfacesdefault high on API & Claude Code · disabling thinking above high returns a 400 | — |
| Latency | Fastest | Fast | Moderate | Slower |
| Tokenizer | Previous gen | New~30% more tokens per text | Newsame generation as 4.8 · ~30% more tokens per text | New~30% more tokens per text |
| Reliable knowledge | Feb 2025 | Jan 2026 | May 2026 Jan 2026the freshest cutoff in the lineup — newer than Fable 5’s | Jan 2026 |
| Launched | Oct 2025 | Jun 30, 2026 | Jul 24, 2026 May 28, 2026 | Jun 9, 2026 |
Cache writes bill at 1.25× base input (5-minute TTL) or 2× (1-hour TTL); reads at 0.1×. Batch and caching discounts stack. Opus 4.8 remains available on every platform — and now has a second job: it’s the default fallback lane when Opus 5’s safety classifiers flag a request in Claude, Claude Code and Cowork (optional on the API). Mythos 5 shares Fable 5’s specs and pricing but is invitation-only for approved organisations via Project Glasswing. Source: Anthropic models overview, What’s new in Claude Opus 5, and pricing docs, Aug 4, 2026.
List price is only the opening bid. Output tokens cost 5× input on every tier, cache reads collapse input cost by 10×, and batch halves everything. Your effective rate is a function of workload shape, not the pricing page — and the Opus refresh added one more lever: a 512-token cache minimum.
The biggest lever nobody budgets for. Agents that re-read a system prompt or codebase pay a tenth of base input on every hit.
Anything latency-tolerant — evals, backfills, report generation — should never pay the interactive rate. Stacks with caching.
Long-TTL writes cost double base input. Pays for itself after two reads — model the hit rate before enabling.
Research preview, Claude API only: $10/$50 for roughly 2.5× the output speed. Latency, like intelligence, is now a line item.
Down from 1,024 on Opus 4.8. Short tool preambles and compact system prompts that never cached before now do — free money for chatty agents.
All four current models bill the full context window at standard per-token rates — a 900k-token request costs the same rate as a 9k one. Web search adds $10 per 1,000 searches on top of tokens.
Capability climbs to the top-right — and you pay for the top-right twice: in dollars and in seconds. The routing question is never “which model is best”, it’s “which point on this canvas does this task deserve”. On July 24 one bubble moved: Opus 5 climbed a capability band without moving on price or latency. That vertical jump is the whole story of this revision.
Capability bands reflect Anthropic’s own positioning language (“close to the frontier intelligence of Fable 5”); latency is Anthropic’s comparative rating. Blended = simple average of standard input and output list price. This is a positioning map, not a benchmark — see below for the launch numbers, and the caveat that comes with them.
Surpasses every other model, at a lower cost per task than 4.8. The Opus tier’s cost per finished task just collapsed.
At max effort — at half the cost per task. Better cost-for-performance than all others at high, xhigh and max effort.
Novel problem-solving, not memorisation — the widest single margin Anthropic claims for this release.
Computer use: outperforms every model at any given cost point, per Anthropic’s effort-curve charts.
End-to-end business tasks at the same cost per task — and even at lowest effort it still tops every other model.
Knowledge work. The caveat Anthropic itself states: still behind Mythos 5 on cybersecurity tasks — deliberately.
All figures as disclosed in Anthropic’s July 24, 2026 launch materials — vendor-stated, effort-dependent, and not yet independently audited. Benchmarks are a prior, not a verdict: the only benchmark that settles routing is your own eval set on your own prompts. Notably absent at launch: a like-for-like SWE-bench Pro figure to line up against the July guide’s chart — hence a claims board this month, not a bar chart with invented axes.
This is the flow I use when designing routing policies: every workload enters at the top, and every “no” you can answer honestly is margin. The dashed rail is the part most teams forget — a cheap model plus a verifier plus an escalation path beats an expensive default almost everywhere. New since July: Anthropic now ships a version of that rail natively — flagged Opus 5 requests fall back to Opus 4.8, and the API grew a server-side fallbacks parameter with a "default" mode.
Sticker prices don’t answer “what will my traffic cost”. Set your workload shape below — the maths runs on Anthropic’s published rates, including the cache-read discount and the new-tokenizer uplift most calculators ignore. The Opus column now prices Opus 5 (identical rates to 4.8 — that’s the point), with an optional fast-mode toggle.
Monthly = 30 days. Assumes a warm cache (reads at 0.1× base input; write premiums and TTL churn excluded), no tool surcharges, list rates on the first-party API. The tokenizer uplift approximates Anthropic’s stated ~30% token increase on new-tokenizer models; real expansion varies by content. The lever this simulator can’t price is effort: Anthropic states Opus 5’s low and medium settings hold strong quality at a fraction of the tokens and latency — which only your own evals can convert into a number. An estimator, not an invoice.
Notes from building model-routing infrastructure — the details that decide whether the ladder saves you money or quietly costs you it. Two carried over from July, six rewritten by the Opus 5 launch.
Whatever your platform falls back to when nobody chooses is your biggest AI cost line. Note that Anthropic’s own “if you’re unsure, start with” answer changed on Jul 24 — the docs now point at Opus 5. Mine is still Sonnet 5 for the daily 80%. Either is defensible; an unreviewed default is not.
Fable 5, Opus 5 and Sonnet 5 use a new tokenizer that yields ~30% more tokens for the same text (Anthropic-stated; Opus 5 stays on 4.8’s generation). Sticker comparisons against Haiku 4.5 or Sonnet 4.6 that ignore it are off by a third.
Opus 5 is the proof at scale: same $5/$25 sticker, a claimed >2× Frontier-Bench score at lower cost per task. Effort is now the bigger lever than model choice — the ladder runs low → max, thinking can’t be disabled above high, and launch partners report similar quality with roughly a quarter fewer tokens at lower settings. Measure cost per finished task, at the effort you’ll actually run.
At 0.1× reads, a warm 200k-token codebase costs $0.10 per pass on Opus 5 instead of $1.00. And the cacheable minimum just dropped to 512 tokens (from 1,024), so short tool preambles finally cache at all. For agent loops, cache architecture matters more than model choice.
Sonnet 5’s $2/$10 window closes August 31, 2026; standard $3/$15 follows. The side-effect nobody prices: the Sonnet→Opus escalation premium falls from 2.5× to 1.67× on input the same day. Escalation literally gets cheaper when intro pricing ends — run your routing evals at both prices now.
Opus 5’s reliable cutoff is May 2026. Fable 5’s is Jan 2026. For freshness-sensitive work — new frameworks, current APIs, recent events — the mid-tier now knows more than the flagship. Route on recency as well as difficulty; it’s a routing signal that didn’t exist in July.
4.8 remains available on every platform, and flagged Opus 5 requests fall back to it by default in Claude, Claude Code and Cowork. Keep it pinned in your policy as the degraded-mode lane — and remember its thinking and effort semantics differ from 5’s, so test the fallback path, not just the happy path.
You can’t reliably predict task difficulty up front, but you can verify outputs cheaply. Cheap-first routing plus a verifier plus escalation-on-failure turns the four-tier ladder into a cost optimiser. Anthropic has now blessed the shape as a platform primitive — server-side fallbacks with a "default" mode — for safety refusals today, but the direction of travel is clear.
The volume floor. If it runs a million times a day and no human reads it raw, it belongs here — escalate the exceptions.
The default. Most production traffic, most coding, most agents — especially while intro pricing holds.
The everyday frontier. Near-Fable intelligence at half the price — the escalation target, and the new daily driver for hard, correctness-critical work.
The long horizon. Reserve it for autonomous, multi-hour-to-multi-day work where quality compounds. A scalpel, not a default — and its edge just got narrower.