On October 7, Anthropic shipped Claude Haiku 5.5: "the cheapest, fastest, and most capable small model" so far. Price and benchmark conditions are covered separately in the Haiku 5.5 pricing post: two-tier pricing at 100K tokens, the new tokenizer, and what Terminal-Bench 39.2% and OSWorld 72.4% measured. Start there for the fine print.
This post answers a different question: where do you actually put this cheap model to work? Anthropic names the slots plainly — classification, summarization, routing, subagents. Not one big model for everything, but small frequent jobs on Haiku and hard judgment escalated upward.
What is different: effort controls come to Haiku
This is the first Haiku with an adjustable effort knob, Low to Max, defaulting to medium.
Low effort (low/medium): cheap and fast, for easy jobs
→ medium default reaches 1277 on GDPval-AA at ~1/10 the max output tokens
High effort (max): costly and slow, for hard jobs
→ 1620 on GDPval-AA, 39.2% on Terminal-Bench, 72.4% OSWorld partial
The official accuracy-vs-cost curves say the same. Haiku 5.5 opens a new cheap region while complex terminal coding stays with Sonnet and Opus — a ~31pp gap on Terminal-Bench (70.6% vs 39.2%). A cheap model forces a placement decision, not a wholesale switch.
The spec is workmanlike. API id claude-haiku-5-5, 1M-token context, 128K max output, June 2026 cutoff, same short id on Bedrock, Vertex, and Foundry. Pricing is $0.10 input and $0.50 output per million up to 100K, times five above ($0.50 and $2.50). The ~75% average saving already reflects the tokenizer change (~30% more tokens), so recount near-100K prompts with the new model.
Translating benchmarks into jobs
The launch table in work language — all Anthropic-measured at max effort, not independent:
| Evaluation | Haiku 5.5 | Haiku 4.5 | Luna (same price) | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (knowledge work) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1 offline partial | 72.4% | 15.7% | 48.9% | 83.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 Main | 46.4% | — | 42.4% | 52.1% |
| HLE no tools / with tools | 45.9% / 57.4% | 10.2% / 18.7% | — | 56.9% / 64.5% |
Where Haiku 5.5 leads Luna (same $0.10/$0.50):
knowledge, computer use, terminal, visual reasoning — all of them
→ at equal price, Haiku claims the edge
Where Sonnet 5.5 still leads:
31pp on Terminal-Bench, 11pp OSWorld partial, 220 Elo points
→ complex coding stays upstairs
Customer quotes agree on placement. Asana reports ~30% lower latency and up to 2.5x per turn, HubSpot 92.8% on a CRM suite, Box up 11 points at half the latency, Rogo extracting 10-K figures with Haiku subagents, Cognition at 66.2 on FrontierCode with Haiku as the Devin Fusion sidekick. All vendor-reported, but the placement matches: short frequent jobs, and subagents beside large models.
Independent notes add caution. Artificial Analysis scores Intelligence Index 43 (up 26 in a year, between Luna at 38 and Sonnet 5.5 at 56) while noting $0.21 per task at max and $0.12 at xhigh, with ~3x Luna's output tokens at max. AutomationBench shows 35% with over-refusal. Cheap is not automatically cheap — read effort and token volume together.
CodeBridge Mini Lab: the 100-job comparison
Straight from the briefing: pull 100 of your current summary and classification jobs and compare.
# volume_lab.py — sketch, structure only
EFFORT = "medium" # start at default, measure before raising
def run_batch(tasks, call_model, log):
for task in tasks[:100]:
# 1. Cheap tier first (Haiku class)
r = call_model("haiku-5-5", task, effort=EFFORT)
ok = verify(r, task)
log.record(model="haiku", cost=r.cost,
latency=r.latency, tokens=r.tokens, ok=ok)
# 2. Escalate only failures (not everything)
if not ok:
r2 = call_model("sonnet-5-5", task, effort="high")
ok2 = verify(r2, task)
log.record(model="sonnet", cost=r2.cost,
latency=r2.latency, tokens=r2.tokens, ok=ok2)
return log.summarize(by=["model"])
Four metrics (per completed job):
[ ] accuracy — same jobs, 3+ runs each, split by effort
[ ] latency — to completion, medium and max separately
[ ] escalation rate — target 10–20% reaching the upper model
[ ] cost per completion — per success, not per token
Three placement rules:
[ ] summaries, compactions, lookups, classification → Haiku by default
[ ] coding subagents (extract, tidy, draft) → Haiku plus verification
[ ] complex terminal and long reasoning → escalate to Sonnet or Opus
Two traps:
[ ] recount near-100K prompts with claude-haiku-5-5
[ ] log effort and trial counts beside Terminal-Bench-style scores
Pair the routing structure with the cost-per-success math and the experiment is complete. Fold in the same-day Sonnet 5.5 cache-read cut ($0.20 → $0.10, ~20% cheaper agentic work). Top and bottom got cheaper together.
Conclusion: skill is where you place the small model
One line to close.
Measure cost per success, not per token — and place Haiku at the front and beside (as subagents).
Haiku 5.5 is the first credibly agentic Haiku. OSWorld 15.7% → 72.4% and Terminal-Bench 0.0% → 39.2% prove it. Gaps to Sonnet, effort-dependent scores, token growth, and the 100K cliff all remain. Read it as a placement table, not a price table, and let a 100-job log set next quarter's model layout.
Further reading
- Why Haiku 5.5 is 90% cheaper: conditions before price
- Model routing in code: cheap models first, escalate when stuck
- Reading AI price tables by cost per success
References
- Anthropic: Introducing Claude Haiku 5.5 (Oct 7, 2026)
- Artificial Analysis: Claude Haiku 5.5 release (Oct 7, 2026)
- Simon Willison: Claude Haiku 5.5 notes (Oct 7, 2026)
Go deeper with a course
To connect cheap-first routing and verification loops in code, this course stacks harness, loop, and graph in order — exactly like the 100-job experiment here.