September 2026 made model comparison messy again. Anthropic shipped Claude Opus 5.5, and OpenAI shipped GPT-6 Astra.
Current Artificial Analysis figures put Opus 5.5 at 58 on the Intelligence Index at max effort, Astra at 53. On numbers alone, Opus 5.5 looks like the answer.
Real selection isn't that simple.
Separate the public data first
Claude Opus 5.5, released September 22, 2026, lowers cost versus Opus 5 while pushing complex coding and knowledge-work performance. Anthropic describes Fable 5.1-level performance on most work at a lower operating cost than Opus 5.
Independent Artificial Analysis figures show this gap:
| Item | Claude Opus 5.5 (max) | GPT-6 Astra (max) |
|---|---|---|
| Intelligence Index | 58 | 53 |
| Input price / 1M tokens | $4 | $10 |
| Output price / 1M tokens | $20 | $50 |
| Cost per Intelligence task | About $5.98 | About $3.26 |
| Context window | 1M | 1M |
Here's the fun part.
Opus 5.5 charges less per token than Astra, yet its measured per-task cost on Intelligence Index work runs higher. The driver: Opus 5.5 at max effort burns far more reasoning and output tokens.
So:
Lower token price
≠
Lower real task cost, always
Where Opus 5.5 looks strong
Artificial Analysis reports Opus 5.5 as very strong across several Intelligence Index sub-evaluations — especially long knowledge work and complex expert tasks like AA-Briefcase, GDPval-AA, and AutomationBench-AA.
Terminal-Bench 4.0 tells a similar story at about 59.6%, roughly matching GPT-6 Astra at xhigh.
So the interesting workloads go beyond simple Q&A:
- Multi-step work across a large codebase
- Reading sources, analyzing, then producing deliverables
- Tool-using work with self-checks
- Tasks holding long context together
Why Astra stays interesting
GPT-6 Astra is OpenAI's September 2026 frontier model. OpenAI positions it on software engineering, browsing, computer use, and expert knowledge work.
Its max score trails Opus 5.5, but Astra's efficiency — strong results from relatively few tokens — stands out. That's exactly why Astra max shows a lower cost per task than Opus 5.5 max under the same evaluation.
So the question changes from "which is smarter?" to this:
Does one deep, top-quality result matter most?
vs
Do you need to repeat many tasks affordably?
CodeBridge Mini Lab: compare with 3 runs each in the same repo
To compare models yourself, skip "which codes better?" Assign the same small real task in the same repo instead.
Prepare a small web project, for example:
Task
1. Fix the validation bug in the login form
2. Keep existing tests green
3. Add 1 new regression test
4. Explain the change in 5 lines or fewer
Give both models identical instructions and run each 3 times.
These observations are plenty:
| Item | How to check |
|---|---|
| Success | Did all tests pass |
| Change scope | Did it touch more files than needed |
| Retries | How many fixes after failure |
| Time | Time to task completion |
| Cost | Total API cost |
| Explanation quality | Can it say why it changed things |
Record results like this:
Model: ______
Run: 1 / 2 / 3
Tests passed: Y / N
Files changed: __
Retries: __
Time: __ sec
Cost: $__
Unexpected changes: Y / N
The key insight: the same model doesn't need to win all 3 runs. Variance across runs often matters more in practice.
You can also split roles between models
You don't have to pick one model at all.
Split it like this, for example:
Fast exploration / drafts
→ Relatively cheap model
Hard analysis / final review
→ Top-performance model
Or:
Astra
→ Fast orientation + driving the work forward
Opus 5.5
→ Hard code review + final checks
The point isn't brand loyalty. It's splitting work into stages.
Related posts
- Why benchmark scores alone mislead you
- Judge AI model cost by cost per successful task
- Why the harness changes results more than the model
References
- Anthropic: Introducing Claude Opus 5.5
- OpenAI: GPT-6 Astra
- Artificial Analysis: Claude Opus 5.5
- Artificial Analysis: GPT-6 Astra
Conclusion: model choice is task design, not ranking
Published totals alone put Opus 5.5 on top. But Astra shows a different trade-off on cost, token use, and some agentic tasks.
So set your selection bar here:
Does one peak-quality result matter, or do repeatable cost and speed matter?
And make the final call with small repeated trials on your repo, your docs, your work — not public benchmarks.
Go deeper with a course
If you want a decision framework for splitting work across models instead of picking sides, a structured course on practical multi-AI workflows helps.