Opus 5.5 landed September 22, and Sonnet 5.5 measurements followed on Artificial Analysis.
The scoreboard reads: Opus 5.5 max at 58, Opus 5.5 xhigh and Sonnet 5.5 max tied at 56.
Which invites an obvious question. Two points apart — why not always run Opus max?
Real usage answers differently. Once you count tokens per task and time alongside price, Sonnet and lower efforts win across a wide range.
Read the scoreboard properly first
Line up the Artificial Analysis Intelligence Index leaders:
| Model | Intelligence Index | Notes |
|---|---|---|
| Claude Opus 5.5 (max) | 58 | Current top score |
| Claude Opus 5.5 (xhigh) | 56 | |
| Claude Sonnet 5.5 (max) | 56 | Tied with Opus xhigh |
| Claude Opus 5.5 (high) | 54 | |
| Claude Fable 5.1 (max) | 53 | Previous-gen leader |
| GPT-6 Astra (max) | 53 | |
| GPT-6.1 Sol (max) | 52 | Joined Sept 29 |
The eye-catcher: Sonnet 5.5 max ties Opus 5.5 xhigh. The lower tier's maximum output matches the upper tier one notch down.
Then September 29–30 added GPT-6.1 Sol (52 points, under a quarter of Astra's task cost) and Gemini 4 Argon (53 points, about $1.99 per task with discounts). The current is obvious. The game isn't one top score anymore — every score band now has its own value seat.
Old choice: pricey top model vs cheap old model (pick 1 of 2)
New choice: 58 / 56 / 54 / 53 / 52-point bands at different prices
→ The question becomes "how many points does this task need?"
Why Opus max really costs more: it burns tokens
Opus 5.5 lists at $4 input and $20 output per 1M tokens. That's 20% below Opus 5, with cache reads down from $0.50 to $0.20 — a 60% cut.
But Artificial Analysis measures about 119,000 output tokens per Intelligence task for Opus 5.5 max. Compare Opus 5 max at about 73,000, Fable 5.1 max at about 78,000, and GPT-6 Astra max at about 27,000. Opus 5.5 max uses 1.5x to 4x more.
The structure looks like this:
Cost of 1 task = token price × tokens used
Opus 5.5 max: lower unit price, but heavy token use at max
→ Head-to-head max comparisons cost more per task than expected
Anthropic's own announcement hints at it. Opus 5.5 at default (medium) beats Opus 5 max on FrontierCode at 54.6% and Terminal-Bench 4.0, matching GPT-6 Astra max at roughly 20–40% of the cost. The story leads with medium, not max — for good reason.
Effort first, tier second
Opus 5.5 offers five levels: low, medium, high, xhigh, max. Artificial Analysis places four of them — max, xhigh, high, medium — on the intelligence-vs-cost Pareto frontier. Plainly: each effort earns its price.
So run your selection in this order:
Step 1: Baseline with Opus max, 3 runs (log success, time, cost)
Step 2: Step the same Opus down high → medium, find where it holds
Step 3: If medium holds, switch to Sonnet max / Sonnet high
Step 4: If that holds, lock in low effort + cache reads
Anthropic's published cases suggest low effort covers more than you'd guess. In Deloitte testing, Opus 5.5 at its lowest setting caught 72% of code-review bugs, and a finance analytics evaluation had the lowest setting beating Opus 5's high setting. Vendor-supplied figures, so don't swallow them whole — but the direction is clear. Start low, climb as needed.
CodeBridge Mini Lab: when a 2-point gap moves money
Pick one small task in your repo and run both sides 3 times each:
Task: fix 1 login-form validation bug + keep existing tests + add 1 regression test
Log every run:
Model / Effort: ______
Tests passed: Y / N
Files changed: __ (touched more than needed?)
Time: __ min
Cost: $__
Explanation quality: 1–5 (is the why readable?)
Fix your grading rule in advance:
- 3 successes out of 3 with readable explanations → adopt that combo
- Sonnet max matching Opus high on success → take the cheaper side
- Any failure at all → check "was the instruction vague?" before raising effort
The classic mistake: deciding after a single run. Agent work wobbles run to run, so 3 runs is the minimum. It's the same method as the Pass@1, cost, and time post.
Related posts
- Claude Opus 5.5 vs GPT-6 Astra: real differences beyond scores
- Judge AI model cost by cost per successful task
- Reasoning high vs max: is thinking longer always better
References
- Anthropic: Introducing Claude Opus 5.5
- Artificial Analysis: Claude Opus 5.5 takes the top spot
- Artificial Analysis: Benchmarking GPT-6 Astra
- Artificial Analysis: Gemini 4 Argon
- Artificial Analysis: GPT-6.1 Sol replaces GPT-6 Sol
Conclusion: ask a different question
Not "Opus or Sonnet?" Ask this instead:
Does this task need 58 points, is 56 enough, or does 54 cover it?
In my experience, doc summaries, test additions, and routine migrations usually clear the bar at Sonnet level or low Opus effort. Long-haul work that loses its way easily — hundred-thousand-line migrations, multi-repo audits — is where high-effort Opus earns its keep. Anthropic showcasing a 680,000-line migration and a 200,000-line audit as Opus 5.5 stories says the same thing.
One line to close: don't default to max — climb up from low effort and lower tiers. Fewer tasks need those 2 points than you'd think.
Go deeper with a course
If you want to build the habit of varying effort and tier in your own repo — measuring success, time, and cost each time — a structured course on real-project harnesses helps.