On October 7, Anthropic shipped Claude Haiku 5.5: "the cheapest, fastest, and most capable small model" so far. Price and benchmark conditions are covered separately in the Haiku 5.5 pricing post: two-tier pricing at 100K tokens, the new tokenizer, and what Terminal-Bench 39.2% and OSWorld 72.4% measured. Start there for the fine print.

This post answers a different question: where do you actually put this cheap model to work? Anthropic names the slots plainly — classification, summarization, routing, subagents. Not one big model for everything, but small frequent jobs on Haiku and hard judgment escalated upward.

What is different: effort controls come to Haiku

This is the first Haiku with an adjustable effort knob, Low to Max, defaulting to medium.

Low effort (low/medium): cheap and fast, for easy jobs
 → medium default reaches 1277 on GDPval-AA at ~1/10 the max output tokens
High effort (max): costly and slow, for hard jobs
 → 1620 on GDPval-AA, 39.2% on Terminal-Bench, 72.4% OSWorld partial

The official accuracy-vs-cost curves say the same. Haiku 5.5 opens a new cheap region while complex terminal coding stays with Sonnet and Opus — a ~31pp gap on Terminal-Bench (70.6% vs 39.2%). A cheap model forces a placement decision, not a wholesale switch.

The spec is workmanlike. API id claude-haiku-5-5, 1M-token context, 128K max output, June 2026 cutoff, same short id on Bedrock, Vertex, and Foundry. Pricing is $0.10 input and $0.50 output per million up to 100K, times five above ($0.50 and $2.50). The ~75% average saving already reflects the tokenizer change (~30% more tokens), so recount near-100K prompts with the new model.

Translating benchmarks into jobs

The launch table in work language — all Anthropic-measured at max effort, not independent:

Evaluation Haiku 5.5 Haiku 4.5 Luna (same price) Sonnet 5.5
GDPval-AA v2.1 (knowledge work) 1620 735 1437 1840
AA-Briefcase v1.1 1578 614 1336 1824
OSWorld 2.1 offline partial 72.4% 15.7% 48.9% 83.9%
Terminal-Bench 4.0 39.2% 0.0% 16.4% 70.6%
FrontierCode 1.1 Main 46.4% — 42.4% 52.1%
HLE no tools / with tools 45.9% / 57.4% 10.2% / 18.7% — 56.9% / 64.5%
Where Haiku 5.5 leads Luna (same $0.10/$0.50):
 knowledge, computer use, terminal, visual reasoning — all of them
 → at equal price, Haiku claims the edge

Where Sonnet 5.5 still leads:
 31pp on Terminal-Bench, 11pp OSWorld partial, 220 Elo points
 → complex coding stays upstairs

Customer quotes agree on placement. Asana reports ~30% lower latency and up to 2.5x per turn, HubSpot 92.8% on a CRM suite, Box up 11 points at half the latency, Rogo extracting 10-K figures with Haiku subagents, Cognition at 66.2 on FrontierCode with Haiku as the Devin Fusion sidekick. All vendor-reported, but the placement matches: short frequent jobs, and subagents beside large models.

Independent notes add caution. Artificial Analysis scores Intelligence Index 43 (up 26 in a year, between Luna at 38 and Sonnet 5.5 at 56) while noting $0.21 per task at max and $0.12 at xhigh, with ~3x Luna's output tokens at max. AutomationBench shows 35% with over-refusal. Cheap is not automatically cheap — read effort and token volume together.

CodeBridge Mini Lab: the 100-job comparison

Straight from the briefing: pull 100 of your current summary and classification jobs and compare.

# volume_lab.py — sketch, structure only
EFFORT = "medium"  # start at default, measure before raising

def run_batch(tasks, call_model, log):
    for task in tasks[:100]:
        # 1. Cheap tier first (Haiku class)
        r = call_model("haiku-5-5", task, effort=EFFORT)
        ok = verify(r, task)
        log.record(model="haiku", cost=r.cost,
                   latency=r.latency, tokens=r.tokens, ok=ok)
        # 2. Escalate only failures (not everything)
        if not ok:
            r2 = call_model("sonnet-5-5", task, effort="high")
            ok2 = verify(r2, task)
            log.record(model="sonnet", cost=r2.cost,
                       latency=r2.latency, tokens=r2.tokens, ok=ok2)
    return log.summarize(by=["model"])
Four metrics (per completed job):
 [ ] accuracy — same jobs, 3+ runs each, split by effort
 [ ] latency — to completion, medium and max separately
 [ ] escalation rate — target 10–20% reaching the upper model
 [ ] cost per completion — per success, not per token

Three placement rules:
 [ ] summaries, compactions, lookups, classification → Haiku by default
 [ ] coding subagents (extract, tidy, draft) → Haiku plus verification
 [ ] complex terminal and long reasoning → escalate to Sonnet or Opus

Two traps:
 [ ] recount near-100K prompts with claude-haiku-5-5
 [ ] log effort and trial counts beside Terminal-Bench-style scores

Pair the routing structure with the cost-per-success math and the experiment is complete. Fold in the same-day Sonnet 5.5 cache-read cut ($0.20 → $0.10, ~20% cheaper agentic work). Top and bottom got cheaper together.

Conclusion: skill is where you place the small model

One line to close.

Measure cost per success, not per token — and place Haiku at the front and beside (as subagents).

Haiku 5.5 is the first credibly agentic Haiku. OSWorld 15.7% → 72.4% and Terminal-Bench 0.0% → 39.2% prove it. Gaps to Sonnet, effort-dependent scores, token growth, and the 100K cliff all remain. Read it as a placement table, not a price table, and let a 100-job log set next quarter's model layout.

Further reading

References

Go deeper with a course

To connect cheap-first routing and verification loops in code, this course stacks harness, loop, and graph in order — exactly like the 100-job experiment here.