STENCEL.AI // MODEL ROUTING FIELD GUIDE
Field guide · Claude model lineup · July 2026

Stop choosing a model.
Write a routing policy.

Anthropic now sells four tiers at once — Haiku 4.5, Sonnet 5, Opus 4.8 and Fable 5 — spanning 10× on input price and 50× end-to-end. Treating that menu as a single choice is how teams overpay. Treating it as a scheduling problem is how platforms win. Verified prices, honest positioning, a decision flow, and a live cost simulator below.

The price ladder · $ per million tokens · log scale · filled = input, ring = output
THE FULL LADDER SPANS 50× — FROM $1 (HAIKU INPUT) TO $50 (FABLE 5 OUTPUT) $1 $2 $3 $5 $10 $25 $50 $ / MTOK · LOG SCALE FABLE 5 $10 $50 OPUS 4.8 $5 $25 SONNET 5 $2 $10 $3/$15 STD FROM SEP 1 HAIKU 4.5 $1 $5

Prices from Anthropic’s official pricing page, retrieved July 4, 2026. Sonnet 5 introductory pricing runs through August 31, 2026. Notice the rhyme: Opus 4.8’s input costs exactly what Haiku 4.5’s output costs — the ladder is deliberately spaced.

01 · The lineup

Four tiers, one API — know what each one is for

Anthropic no longer ships a single flagship. It ships a price-segmented ladder, and every tier has a legitimate job. The one most comparisons skip — Haiku — is the one that should be carrying most of your call volume.

Volume tier

Haiku 4.5

claude-haiku-4-5
$1 / $5 per MTok in / out · batch $0.50 / $2.50

The fastest Claude, with near-frontier intelligence. The only current model on the previous tokenizer and a 200k window — and priced for anything you run millions of times a day.

Context200k
Max output64k
LatencyFastest
ThinkingExtended
Reach for it whenClassification, triage, extraction, guardrail checks — latency-critical, high-volume calls where unit cost dominates.
Production default

Sonnet 5

claude-sonnet-5
$2 / $10 intro per MTok · $3 / $15 from Sep 1, 2026
Intro pricing ends Aug 31

The best balance of speed and intelligence in the lineup — the most agentic Sonnet yet, built to plan and drive tools autonomously. Launched June 30, 2026.

Context1M
Max output128k
LatencyFast
ThinkingAdaptive
Reach for it whenDay-to-day coding, agents and document work — where the majority of production traffic should land.
Correctness tier

Opus 4.8

claude-opus-4-8
$5 / $25 per MTok · fast mode (preview) $10 / $50

Built for complex agentic coding and enterprise work. Effort defaults to high on every surface, and a research-preview fast mode roughly doubles the rate when latency matters. Launched May 28, 2026.

Context1M
Max output128k
LatencyModerate
ThinkingAdaptive
Reach for it whenCorrectness-critical refactors and hard multi-step reasoning — work where a wrong answer is expensive.
Frontier tier

Fable 5

claude-fable-5
$10 / $50 per MTok · batch $5 / $25

Next-generation intelligence for long-running agents — the generally available half of the Mythos-class tier, shipped with dual-use safety classifiers (Mythos 5 is the invitation-only variant). Adaptive thinking is always on. GA June 9, 2026.

Context1M
Max output128k
LatencySlower
ThinkingAlways-on
Reach for it whenLong-horizon autonomous agents where task quality dominates token cost.
02 · Head to head

The full spec & pricing sheet

Everything below is from Anthropic’s official model and pricing documentation. The rows most people skip — cache-read rates and the tokenizer generation — are the ones that decide real-world cost.

Spec Haiku 4.5 Sonnet 5 Opus 4.8 Fable 5
PositioningFastest, near-frontier intelligenceBest speed-to-intelligence balanceComplex agentic coding & enterprise workNext-gen intelligence for long-running agents
Input / output$1 / $5$2 / $10intro to Aug 31 · then $3 / $15$5 / $25fast mode $10 / $50$10 / $50
Cache read (0.1×)$0.10$0.20then $0.30$0.50$1.00
Batch (50% off)$0.50 / $2.50$1 / $5then $1.50 / $7.50$2.50 / $12.50$5 / $25
Context window200k1Mstandard pricing across full window1Mstandard pricing across full window1Mstandard pricing across full window
Max output64k128k300k via Batch API beta128k300k via Batch API beta128k
ThinkingExtendedAdaptiveAdaptiveAdaptive · always on
Effort defaultHighon API & Claude CodeHighon all surfaces
LatencyFastestFastModerateSlower
TokenizerPrevious genNew~30% more tokens per textNew~30% more tokens per textNew~30% more tokens per text
Reliable knowledgeFeb 2025Jan 2026Jan 2026Jan 2026
LaunchedOct 2025Jun 30, 2026May 28, 2026Jun 9, 2026

Cache writes bill at 1.25× base input (5-minute TTL) or 2× (1-hour TTL); reads at 0.1×. Batch and caching discounts stack. Mythos 5 shares Fable 5’s specs and pricing but is invitation-only for approved organisations. Source: Anthropic models overview & pricing docs, July 4, 2026.

03 · Price geometry

Where the money actually goes

List price is only the opening bid. Output tokens cost 5× input on every tier, cache reads collapse input cost by 10×, and batch halves everything. Your effective rate is a function of workload shape, not the pricing page.

List price per million tokens

LINEAR SCALE · STANDARD RATES · SONNET SHOWN AT INTRO
$10 $20 $30 $40 $50 $0 $ / MTOK HAIKU 4.5 SONNET 5 OPUS 4.8 FABLE 5 IN $1 OUT $5 IN $2 OUT $10 · std $3/$15 IN $5 OUT $25 IN $10 OUT $50 LIGHT = INPUT · SOLID = OUTPUT
Cache read
0.1×
The biggest lever nobody budgets for. Agents that re-read a system prompt or codebase pay a tenth of base input on every hit.
Batch API
0.5×
Anything latency-tolerant — evals, backfills, report generation — should never pay the interactive rate. Stacks with caching.
1-hour cache write
Long-TTL writes cost double base input. Pays for itself after two reads — model the hit rate before enabling.
Fast mode · Opus 4.8
Research preview: $10/$50 for materially faster Opus output. Latency, like intelligence, is now a line item.

All four current models bill the full context window at standard per-token rates — a 900k-token request costs the same rate as a 9k one. Web search adds $10 per 1,000 searches on top of tokens.

04 · Capability × latency × cost

The trade-off, on one canvas

Capability climbs to the top-right — and you pay for the top-right twice: in dollars and in seconds. The routing question is never “which model is best”, it’s “which point on this canvas does this task deserve”.

Positioning map

BUBBLE AREA ∝ BLENDED LIST PRICE (STANDARD)
Near-frontier Frontier Frontier+ Next-gen FASTEST FAST MODERATE SLOWER COMPARATIVE LATENCY, PER ANTHROPIC HAIKU 4.5 $3 blended SONNET 5 $9 blended OPUS 4.8 $15 blended FABLE 5 $30 blended

Capability tiers reflect Anthropic’s own positioning language; latency is Anthropic’s comparative rating. Blended = simple average of standard input and output list price. This is a positioning map, not a benchmark — see below for one of those.

SWE-bench Pro · agentic coding

ANTHROPIC-REPORTED FIGURES · DIRECTIONAL
FABLE 5 OPUS 4.8 SONNET 5 SONNET 4.6 80.3%vendor-stated 69.2% 63.2% 58.1%

Scores as disclosed around the May–June 2026 launches (Fable 5 figure is Anthropic-stated, not independently audited; Haiku 4.5 is not reported on this benchmark). Benchmarks are a prior, not a verdict — the only benchmark that settles routing is your own eval set on your own prompts.

05 · The decision flow

Route cheap-first, escalate on verified failure

This is the flow I use when designing routing policies: every workload enters at the top, and every “no” you can answer honestly is margin. The dashed rail is the part most teams forget — a cheap model plus a verifier plus an escalation path beats an expensive default almost everywhere.

NEW WORKLOAD Millions of calls? Latency-critical? Simple shape — classify, triage, extract, guardrail-check? YES NO Haiku 4.5 $1/$5 · THE VOLUME FLOOR Standard production work? Coding, agents, document processing, analysis — the daily 80%? YES NO Sonnet 5 $2/$10 INTRO · THE DEFAULT Correctness-critical? Hard multi-step reasoning, refactors that must not regress, expensive to get wrong? YES NO Opus 4.8 $5/$25 · THE CORRECTNESS TIER Long-horizon autonomous agents Multi-hour or multi-day runs where task quality dominates token cost THEN Fable 5 $10/$50 · THE LONG HORIZON ESCALATE ON VERIFIED FAILURE Route cheap-first · verify the output · escalate only what fails. Every hop you avoid is margin.
06 · The simulator

Price your own workload, live

Sticker prices don’t answer “what will my traffic cost”. Set your workload shape below — the maths runs on Anthropic’s published rates, including the cache-read discount and the new-tokenizer uplift most calculators ignore.

HAIKU 4.5
SONNET 5
OPUS 4.8
FABLE 5

Monthly = 30 days. Assumes a warm cache (reads at 0.1× base input; write premiums and TTL churn excluded), no tool surcharges, list rates on the first-party API. The tokenizer uplift approximates Anthropic’s stated ~30% token increase on new-tokenizer models; real expansion varies by content. An estimator, not an invoice.

07 · Field notes

Six things the pricing page won’t tell you

Notes from building model-routing infrastructure — the details that decide whether the ladder saves you money or quietly costs you it.

01

Your default model is a budget decision

Whatever your platform falls back to when nobody chooses is your biggest AI cost line. Make the default an explicit, reviewed policy — Sonnet 5 today — not an accident of whoever wrote the first config.

02

Mind the tokenizer tax

Fable 5, Opus 4.8 and Sonnet 5 use a new tokenizer that yields ~30% more tokens for the same text (Anthropic-stated). Sticker comparisons against Haiku 4.5 or Sonnet 4.6 that ignore it are off by a third — Anthropic itself calls Sonnet 5’s intro price roughly cost-neutral versus 4.6 for exactly this reason.

03

Cheaper per token ≠ cheaper per task

Independent lab Artificial Analysis measured Sonnet 5 at ~$2.29 per task versus ~$1.99 for Opus 4.8 on its Intelligence Index at standard rates — the cheaper model spent more tokens and turns. Effort settings swing turn counts up to ~6×, a bigger cost lever than model choice. Measure cost per finished task.

04

Caching is the real discount ladder

At 0.1× reads, a warm 200k-token codebase costs $0.10 per pass on Opus 4.8 instead of $1.00. Structure prompts so the stable prefix caches — for agent loops, cache architecture matters more than model choice.

05

Diary date: August 31

Sonnet 5’s $2/$10 window closes August 31, 2026; standard $3/$15 follows. Run your routing evals at both prices now so September’s invoice isn’t a surprise — and remember batch halves whichever rate applies.

06

Escalation beats prediction

You can’t reliably predict task difficulty up front, but you can verify outputs cheaply. Cheap-first routing plus a verifier plus escalation-on-failure turns the four-tier ladder into a cost optimiser — this is how I build routers, and it beats any single default.

08 · Verdicts

If you only remember four lines

HAIKU 4.5

The volume floor. If it runs a million times a day and no human reads it raw, it belongs here — escalate the exceptions.

SONNET 5

The default. Most production traffic, most coding, most agents — especially while intro pricing holds.

OPUS 4.8

The correctness tier. When wrong is expensive, the premium over Sonnet is the cheapest insurance you can buy.

FABLE 5

The long horizon. Reserve it for autonomous, multi-hour work where quality compounds. A scalpel, not a default.