September 2026 made model comparison messy again. Anthropic shipped Claude Opus 5.5, and OpenAI shipped GPT-6 Astra.

Current Artificial Analysis figures put Opus 5.5 at 58 on the Intelligence Index at max effort, Astra at 53. On numbers alone, Opus 5.5 looks like the answer.

Real selection isn't that simple.

Separate the public data first

Claude Opus 5.5, released September 22, 2026, lowers cost versus Opus 5 while pushing complex coding and knowledge-work performance. Anthropic describes Fable 5.1-level performance on most work at a lower operating cost than Opus 5.

Independent Artificial Analysis figures show this gap:

Item Claude Opus 5.5 (max) GPT-6 Astra (max)
Intelligence Index 58 53
Input price / 1M tokens $4 $10
Output price / 1M tokens $20 $50
Cost per Intelligence task About $5.98 About $3.26
Context window 1M 1M

Here's the fun part.

Opus 5.5 charges less per token than Astra, yet its measured per-task cost on Intelligence Index work runs higher. The driver: Opus 5.5 at max effort burns far more reasoning and output tokens.

So:

Lower token price
≠
Lower real task cost, always

Where Opus 5.5 looks strong

Artificial Analysis reports Opus 5.5 as very strong across several Intelligence Index sub-evaluations — especially long knowledge work and complex expert tasks like AA-Briefcase, GDPval-AA, and AutomationBench-AA.

Terminal-Bench 4.0 tells a similar story at about 59.6%, roughly matching GPT-6 Astra at xhigh.

So the interesting workloads go beyond simple Q&A:

  • Multi-step work across a large codebase
  • Reading sources, analyzing, then producing deliverables
  • Tool-using work with self-checks
  • Tasks holding long context together

Why Astra stays interesting

GPT-6 Astra is OpenAI's September 2026 frontier model. OpenAI positions it on software engineering, browsing, computer use, and expert knowledge work.

Its max score trails Opus 5.5, but Astra's efficiency — strong results from relatively few tokens — stands out. That's exactly why Astra max shows a lower cost per task than Opus 5.5 max under the same evaluation.

So the question changes from "which is smarter?" to this:

Does one deep, top-quality result matter most?

vs

Do you need to repeat many tasks affordably?

CodeBridge Mini Lab: compare with 3 runs each in the same repo

To compare models yourself, skip "which codes better?" Assign the same small real task in the same repo instead.

Prepare a small web project, for example:

Task
1. Fix the validation bug in the login form
2. Keep existing tests green
3. Add 1 new regression test
4. Explain the change in 5 lines or fewer

Give both models identical instructions and run each 3 times.

These observations are plenty:

Item How to check
Success Did all tests pass
Change scope Did it touch more files than needed
Retries How many fixes after failure
Time Time to task completion
Cost Total API cost
Explanation quality Can it say why it changed things

Record results like this:

Model: ______
Run: 1 / 2 / 3
Tests passed: Y / N
Files changed: __
Retries: __
Time: __ sec
Cost: $__
Unexpected changes: Y / N

The key insight: the same model doesn't need to win all 3 runs. Variance across runs often matters more in practice.

You can also split roles between models

You don't have to pick one model at all.

Split it like this, for example:

Fast exploration / drafts
→ Relatively cheap model

Hard analysis / final review
→ Top-performance model

Or:

Astra
→ Fast orientation + driving the work forward

Opus 5.5
→ Hard code review + final checks

The point isn't brand loyalty. It's splitting work into stages.

Related posts

References

Conclusion: model choice is task design, not ranking

Published totals alone put Opus 5.5 on top. But Astra shows a different trade-off on cost, token use, and some agentic tasks.

So set your selection bar here:

Does one peak-quality result matter, or do repeatable cost and speed matter?

And make the final call with small repeated trials on your repo, your docs, your work — not public benchmarks.

Go deeper with a course

If you want a decision framework for splitting work across models instead of picking sides, a structured course on practical multi-AI workflows helps.