Anthropic released Claude Opus 5.5 on September 22, 2026. It opens the new Claude 5.5 family, and its message is clear: not a smarter model at a new price, but near-top-tier performance at a lower operating cost.

This post organizes the published price and performance numbers, then reads the announcement from a coding-agent user's perspective.

What was announced

Anthropic's pitch compresses to one line:

Fable 5.1-level performance on most work, at 40% lower operating cost than Opus 5.

The price table changed like this:

Item Opus 5 Opus 5.5
Input / 1M tokens $5 $4
Output / 1M tokens $25 $20
Cache reads $0.50 $0.20
Cache writes $6.25 $5

Unit prices fell 20%, yet Anthropic claims 40% lower operating cost — because the model uses fewer tokens per task and generates output over 30% faster. Same work, fewer tokens, finished sooner. Claude Code and the Claude Platform also gained a Fast mode up to 2.5x quicker, priced separately at $8 input and $40 output.

The headline numbers sit in agentic coding:

Terminal-Bench 4.0
Opus 5.5 66.4% vs Fable 5.1 55.8%

FrontierCode
Opus 5.5 54.4% vs Fable 5.1 50.3%

Outside comparisons add color. At default effort on FrontierCode, Opus 5.5 beat GPT-6 Astra at roughly a fifth of the per-task cost, and beat GPT-5.6 Sol on CursorBench by 11 points at about a third of the cost. That's the backdrop behind reports of Opus 5.5 topping GPT-5.6 Sol on some dev evaluations. Anthropic itself adds an honest caveat: at this level, a few points don't translate straight into felt differences. Fair warning, I think.

More memorable than benchmarks are the long-horizon stories. One tester migrated 680,000 lines of code within a day; another ran 18+ hours across six repos on its own. Sixteen of eighteen internal research reports cleared the quality bar, where Opus 5 and Fable 5.1 cleared zero. For an agent-model launch, stories like these matter more than scores.

External review came from Frontier Design and METR, and Anthropic's automated behavior audit scored its best result to date. Sonnet 5.5 and Haiku 5.5 follow within weeks, and the model is available in GitHub Copilot too.

In agents, loop count sets the bill — not unit price

A coding agent never finishes a task in one model call. It loops: read, edit, build, fail, fix again.

Cost of 1 task
=
call count × tokens per call × token price
+
build + test + human review cost

Which gives you this relationship:

20% cheaper token price
≠
20% cheaper task cost, always

Halve the call count
→
Task cost roughly halves at equal unit price

GitHub Copilot's early testing says the same thing: similar task success to Opus 5 with "significantly fewer steps and tokens." If that holds, shorter loops move your invoice more than the price cut.

Read the reverse too. A low-looking unit price can still cost more per task under settings like max effort that burn heavy reasoning tokens. As the Opus 5.5 vs GPT-6 Astra comparison showed, unit price and task cost are different numbers. Judge a model by average calls and tokens on your work — not the price page.

Read the safeguards story as an architecture problem

Opus 5.5 ships with Fable 5.1-class safety classifiers. In high-risk zones — cybersecurity, biology, frontier AI development — flagged requests quietly route to older models. The known mapping looks roughly like this:

Cybersecurity flag → Opus 4.8
Biology / frontier LLM development flag → Opus 5

Anthropic says direct debugging of your own code stays on Opus 5.5. Still, workflow builders should note one thing: across a multi-turn task, some intermediate request may be served by a different model. If your eval assumes every request hit the same model, expect unreproducible wobble.

CodeBridge Mini Lab: measure just 10–20 runs in your own repo

If you use Claude Code or a Codex-class tool, this table beats any public score. Pick 10–20 real tasks in your repo and record three things: success rate, average cost, completion time.

Model: ______
Run: 1 / 2 / 3
Tests passed: Y / N
Files changed: __
Retries: __
Time: __ sec
Cost: $__
Unexpected changes: Y / N

A small CSV works fine:

task_id,model,cost,success,retries,time_sec
1,opus-5-5,0.12,1,0,24
2,opus-5-5,0.31,0,2,88
3,opus-5-5,0.14,1,1,41
import pandas as pd

df = pd.read_csv("runs.csv")
summary = df.groupby("model").agg(
    total_cost=("cost", "sum"),
    successes=("success", "sum"),
    avg_time=("time_sec", "mean"),
    avg_retries=("retries", "mean"),
)
summary["cost_per_success"] = summary["total_cost"] / summary["successes"]
print(summary)

Watch cost_per_success, not unit price. A cheap model that fails three times isn't cheap — and Opus 5.5 is simply the occasion to redo that math.

Related posts

References

Conclusion: model choice is task design, not ranking

Opus 5.5 asks a different question than "is this model number one?" It asks: "what does it cost on average to carry my task to success?" The 20% price cut is the starting point; shorter loops and cheaper cache reads complete the number. Subscriptions point the same way: 20% higher 5-hour limits, lasting about 25% longer thanks to lower costs.

Confirm the final call with small repeated trials in your own repo — not public benchmarks.

Go deeper with a course

If you want to stop running agent coding on gut feel and control it with verification and constraints, a structured course builds the full workflow.