On October 7, Anthropic shipped Claude Haiku 5.5. The headline is loud: 90% lower token prices than Haiku 4.5, about 75% lower average cost per workload.

Read the full announcement and the discount is not the whole story. The first Haiku with adjustable effort, per-benchmark measurement conditions, and a price cliff at 100K tokens all arrive together. This post reads the conditions before the price tag.

Pricing: 90% only holds at or under 100K tokens

The official table, per million tokens.

Item Haiku 5.5 (up to / over 100K) Haiku 4.5 Sonnet 5.5
Input $0.10 / $0.50 $1.00 $2.00
Output $0.50 / $2.50 $5.00 $10.00
Cache reads $0.01 / $0.05 $0.10 $0.10
Cache writes $0.125 / $0.625 $1.25 $2.50

In short:

Prompts up to 100K tokens:  90% below Haiku 4.5
Prompts over 100K tokens:   50% below Haiku 4.5
~90% of Haiku 4.5 requests fell in the first bucket (Anthropic's count)
→ ~75% lower average workload cost (tokenizer change included)

One trap hides here. Haiku 5.5 uses a new tokenizer from the Sonnet and Opus 5.5 family, producing roughly 30% more tokens on the same text. The 75% figure already accounts for that. Sticker price says 90%, realized says 75%.

The sharper trap is the cliff. The whole request pays the tier matching its prompt length. A 100,000-token input costs about $0.010; at 100,001 tokens the entire prompt pays the higher rate, about $0.050. Input times five, output times five. A workload at 90K tokens on Haiku 4.5 can re-tokenize near 117K on Haiku 5.5, cross the line, and land closer to 35% savings than 90%. That is why the migration guide says to recount prompts near 100K with model: claude-haiku-5-5.

Performance: numbers up, gaps intact

The launch benchmark table, all Anthropic-measured at max effort.

Evaluation Haiku 5.5 Haiku 4.5 GPT-6 Luna Sonnet 5.5
Terminal-Bench 4.0 (agentic coding) 39.2% 0.0% 16.4% 70.6%
OSWorld 2.1 offline subset (computer use) 72.4% 15.7% 48.9% 83.9%
FrontierCode 1.1 Main 46.4% — 42.4% 52.1%
GDPval-AA v2.1 (knowledge work) 1620 735 1437 1840
AA-Briefcase v1.1 1578 614 1336 1824
HLE (no tools / with tools) 45.9% / 57.4% 10.2% / 18.7% — 56.9% / 64.5%

As a sketch:

Cheap tier (Haiku 5.5 ahead of Luna):
  Terminal-Bench  39.2% vs 16.4%
  OSWorld subset  72.4% vs 48.9%
  GDPval-AA       1620 vs 1437

Top tier (Sonnet 5.5 still ahead):
  Terminal-Bench  70.6% vs 39.2% (~31pp gap)
  OSWorld subset  83.9% vs 72.4%
  GDPval-AA       1840 vs 1620

The official accuracy-vs-cost charts say the same. Haiku 5.5 opens a new cheap region, while complex terminal coding stays with Sonnet and Opus. Anthropic names the uses plainly: high-volume repetitive work like summaries, compactions, queries, and classification, plus subagents beside bigger models.

Conditions: same score, different effort and partial credit

This is the re-verification core. Numbers without conditions mislead.

Terminal-Bench 4.0 (39.2%):
 - 66 tasks × 10 trials each
 - Claude Code bare mode, max effort
 - safeguards on, no fallback (1.8% blocked and failed)
 - no internet, pre-cached resources (could lower scores)
 - standard error ~±1.9

OSWorld 2.1 (72.4%):
 - offline subset: 82 of 108 tasks
 - offline VM, 1080p, max 500 steps
 - partial-credit basis; strict pass rate is 37.1%
 - Sonnet 5.5 is 83.9% partial / 48.8% strict (read both)
 - Luna's 48.9% was run by Anthropic via OpenAI's API

And the effort trap. The first-ever Haiku effort knob runs Low to Max with medium as default. Terminal-Bench lands near 39% at max and near 20% at medium. Memorizing "Haiku 5.5 scores 39" without that misses why default settings show about half. The rule from the benchmark reading guide and the version-change post holds: always record version, harness, effort, and trial count.

What came along: Sonnet cache cuts and credits

Haiku was not the only change.

1. Sonnet 5.5 cache reads $0.20 → $0.10 (50% cut)
   → ~20% cheaper on most agentic work
2. Monthly API credits for Max and Team subscribers
   → Max 5x $100, Max 20x $200, Team pooled $500
3. Computer use and browser use in Python and TS SDKs (beta)
   → Haiku 5.5 fits that slot on speed and price

That is why the Terminal-Bench chart draws Sonnet at both old and new cache prices. Top models got cheaper to operate too. Routing ingredients improved top and bottom at once.

Customer quotes in the launch point the same way. Asana reports ~30% lower task latency and up to 2.5x per turn, HubSpot 92.8% on a CRM suite, AlphaSense 0.84 vs 0.76 on document queries, Box 11 points up at half the latency, Rogo pulling 10-K lines with Haiku subagents, Cognition holding a 66.2 FrontierCode score with Haiku as the Devin Fusion sidekick. All vendor-reported, not independent, but the placement hint is clear: short frequent jobs, and subagents beside large models.

CodeBridge Mini Lab: run the same job on two models

# cost_lab.py — sketch, structure only
MAX_ESCALATIONS = 2

def run_batch(tasks, call_model, log):
    for task in tasks:
        # 1. Cheap model first (Haiku 5.5 class, medium effort)
        result = call_model("light", task)
        ok = verify(result, task)
        log.record(model="light", cost=result.cost,
                   latency=result.latency, retries=0, ok=ok)
        # 2. Escalate on failure (max 2)
        attempts = 0
        while not ok and attempts < MAX_ESCALATIONS:
            attempts += 1
            result = call_model("strong", task)
            ok = verify(result, task)
            log.record(model="strong", cost=result.cost,
                       latency=result.latency, retries=attempts, ok=ok)
    # 3. Judge on four metrics, not sticker price
    return log.summarize(by=["model"])
The 4-metric set (straight from the briefing):
 [ ] success rate (same job, 3+ runs each)
 [ ] latency (to completion, split by effort)
 [ ] retry count (target escalation rate 10–20%)
 [ ] cost per completed job (per success, not per token)

Two cautions:
 [ ] recount near-100K prompts with the new tokenizer
 [ ] log effort and trial counts beside Terminal-Bench-style scores

Pair the cost-per-successful-task math with the routing structure and the experiment is complete.

Conclusion: a cheap model changes the design

One line to close.

Measure cost per success, not cost per token. Haiku 5.5 lays a new bottom rung for the router.

The 90% cut unpacked: 90% under 100K, 50% above, ~75% average after the tokenizer. Scores ahead of Luna in the cheap tier, a standing gap to top models, and a condition sheet of effort and partial credit. The action after reading all of it is one: run small and large models on the same jobs and log four metrics. That log decides next quarter's model placement.

Further reading

References

Go deeper with a course

To practice designing cost and success rate together through classification, delegation, and verification, this course stacks harness, loop, and graph in order, exactly like the routing experiment here.