First, a confession. Two days ago in my Sonnet vs Opus post, I wrote "start with lower effort and the smaller tier, then scale up." The direction was right. But the detailed Sonnet 5.5 measurements from September 28 force one correction. Sonnet max is far more expensive than I expected.
Artificial Analysis measured Sonnet 5.5 deeply. The result fits in one line. The same model delivered the Terminal-Bench top score and the highest token usage ever measured.
Start with these 4 new numbers
| Item | Sonnet 5.5 max | Comparison |
|---|---|---|
| Intelligence Index | 56 | Tied with Opus xhigh, 2nd overall |
| Output tokens per task | ~193k | Highest ever measured, ~7x Astra max |
| Cost per task | ~$7.60 | ~50% higher than Sonnet 5 |
| Terminal-Bench 4.0 | 64% | No. 1, ahead of Opus 5.5 at 60% and Astra |
The list price did not change. Input is $2 and output is $10 per million tokens, the same as Sonnet 5. Cost per task still jumped because the model uses far more tokens. In the last post I said Opus max uses a lot of tokens. Sonnet max uses about 60% more than that.
Unit price is flat, but cost per task rose 50%
Cause: the same score is pushed through with token volume
→ a trap you miss if you read only the price table
Yet it ranks No. 1 on Terminal-Bench
That is the interesting part. Sonnet 5.5 max scored 64% on Terminal-Bench 4.0. That beats Opus 5.5 at 60% and Astra at 59%. It also nearly matches Opus on AA-Briefcase (1811 vs 1822 Elo), GDPval-AA (1844 vs 1846 Elo), and AutomationBench-AA (71% vs 70%).
So Sonnet max is a strange mix: "Opus-level intelligence, worst-in-class fuel economy." Why use it? Lower the effort and the answer appears.
The real answer: high effort is the sweet spot
The AA analysis finds Sonnet 5.5's high effort setting the most competitive. It delivers near GPT-6 Sol-level intelligence at nearly the same cost per task. Max and xhigh sit outside the Pareto frontier.
Sonnet 5.5 by effort:
- max: best intelligence, worst cost (special cases only)
- high: best value (candidate default)
- medium and below: Sol models use fewer tokens for the same money
The weakness is also clear. Factual knowledge (AA-Omniscience 54% vs Opus 66%) and science reasoning (HLE and SciCode trail by about 6 points) show the smaller tier's limits. Terminal work and office automation reach Opus level. Factual recall sits below Opus. Pick by use case.
Where Sonnet high fits:
- Coding and ops work that runs in the terminal
- Reading documents and building tables and decks
- Repetitive workflows with many tool calls
Where Opus is still worth it:
- Research where factual accuracy is critical
- Problems mixing science and math reasoning
- Long migrations across hundreds of thousands of lines
CodeBridge Mini Lab: add one line to the last experiment
Reuse the 3-run comparison from the last post. Just change the candidates:
Candidates (last time Opus max/high/medium → this time):
1. Sonnet high (new default candidate)
2. Sonnet max (Terminal-Bench-style tasks only)
3. GPT-6.1 Sol max (to check the cost floor)
Same scorecard: success rate, time, cost, explanation quality
Rule: if Sonnet max does not beat high on success rate, pick high
Note that the AA numbers come from a pre-release build. Anthropic says a structured-output bug was fixed in the public release. Scores should hold or rise slightly. As our benchmark-reading guide says, always write down the measurement conditions. That habit pays off here.
Conclusion: here is my updated advice
Last post: "Start with lower effort and the smaller tier."
Updated: "For Sonnet, start with high, not max. Save max for moments when you need that Terminal-Bench No. 1."
The line "few tasks hinge on a 2-point gap" still holds. What is new is how those 2 points are bought: with token volume. Always write token counts next to intelligence scores. As long as you judge by cost per successful task, that habit saves money.
Related posts
- Claude Sonnet 5.5 vs Opus 5.5: max is often not the answer
- Judge AI models by cost per successful task
- Terminal-Bench 4.0 and the new knowledge-work benchmarks
References
- Artificial Analysis: Claude Sonnet 5.5 reaches #2
- Anthropic: Introducing Claude Opus 5.5
- Artificial Analysis: Models Leaderboard
Go deeper with a course
If you want to practice comparing effort levels and tiers by success rate, time, and cost in a real repository, learn with Claude Code and verification loops.