Coding AI benchmark tables changed a lot in September.
I already wrote a SWE-bench, Terminal-Bench, and ProgramBench comparison. That older baseline no longer reads today's scores well. Terminal-Bench 4.0, AA-Briefcase v1.1, and GDPval-AA v2.1 are now mainstream.
This post organizes the new scoreboard as of October 1. It also covers what matters more than scores: the same model gets different numbers depending on who measures it.
Terminal-Bench 4.0: finish the job in the terminal
Terminal-Bench does not grade chat answers. It runs commands in a terminal and checks whether the task is finished. Key scores on version 4.0 look like this:
| Model | Score | Conditions |
|---|---|---|
| Claude Opus 5.5 | 66.4% | Anthropic internal, xhigh, Claude Code harness |
| Claude Opus 5.5 | 59.6% | Artificial Analysis independent test |
| GPT-6 Astra | 59% | Artificial Analysis test |
| GPT-6 Astra | 57.9% | OpenAI report, high effort |
| Claude Opus 5 | 52.3% | Reproduction test |
The two Opus 5.5 numbers, 66.4% and 59.6%, are not a typo. Different harnesses and settings produce different results. Anthropic also reports a standard error of plus or minus 2.6 points.
Lesson 1: every score travels with "who measured, with which harness, how many runs"
Memorizing numbers alone misleads you. Write down the conditions too.
Read these 3 coding benchmarks together
Tests that grade code changes themselves were updated too.
| Test | What it checks | Key scores |
|---|---|---|
| FrontierCode v1.1 | Is the code change merge-worthy | Opus 5.5 54.4%, Astra 53.3%, GPT-5.6 Sol 47.5% |
| CursorBench 4.0 | Ambiguous real-world multi-file work | Opus 5.5 57.8%, Fable 5.1 51.8%, Opus 5 46.6% |
| Terminal-Bench-Science 0.1 | Terminal-based science research workflows | Astra 64.6%, Opus 5.5 58.7%, Opus 5 29.0% |
Notice that Astra beats Opus 5.5 on Terminal-Bench-Science. The overall intelligence leader does not win every test. Each test measures something different, so read the test that resembles your work.
Knowledge-work benchmarks: AA-Briefcase and GDPval
Tests for docs, analysis, and decks have new versions too.
| Test | What it checks | Key scores |
|---|---|---|
| AA-Briefcase v1.1 | Long order-level knowledge work, analysis plus deck quality | Opus 5.5 Elo 1822, +143 over Fable |
| GDPval-AA v2.1 | 44 real professional tasks | Opus 5.5 Elo 1846, Fable 1735, Opus 5 1708 |
| AutomationBench | Cross-app office automation | Opus 5.5 40.0% (Zapier test), Astra 69% (AA implementation) |
| GDP.pdf | Expert document reasoning, full pass rate | Astra 31%, GPT-5.6 Sol 27% |
There is a trap here too. Opus at 40% and Astra at 69% on AutomationBench come from different implementers. The Opus number is Zapier's own eval. The Astra number is an Artificial Analysis implementation. The same test name with a different implementation can behave like a different test.
Lesson 2: the same name with a different implementation is a different test
Do not compare "AutomationBench 69% vs 40%" directly.
Also note that Opus 5.5 is the first Anthropic model to win on both analysis and deck quality in AA-Briefcase, while Gemini 4 Argon topped the rubric pass rate at 65% but scored lower on presentation. In long work, analysis skill and presentation skill can diverge.
A 4-item checklist for reading scores
When you see a new scoreboard, write these four items next to every number:
[ ] Bench version: 4.0, v1.1, or v2.1?
[ ] Harness: Claude Code, Codex, or a reference harness like Stirrup?
[ ] Effort: which of low / medium / high / xhigh / max?
[ ] Run count: single run or 3-5 run average (is there a standard error)?
"Opus 5.5 Terminal-Bench 66.4%" alone is half the story. Add "Anthropic internal, xhigh, Claude Code harness, standard error plus-minus 2.6" so your future self is not confused. This matches the principles in how to read bench scores and why scores change.
CodeBridge Mini Lab: make 5 questions for your own repo
Public benchmarks are someone else's exam. What you really need is 5 questions from your own repo:
1. One recurring bug fix (judged by tests)
2. One refactor across 3+ files (judged by tests plus diff scope)
3. Summarize one doc into one table (a human judges in 5 minutes)
4. Add one small feature with no new dependencies (judged by tests)
5. Explain one failing command (a human judges)
Run each model and effort 3 times, record success rate, time, and cost
→ if the public No. 1 differs from your No. 1, trust yours
This follows the same idea as rebuilding ProgramBench from scratch. You shrink the rebuild-everything test into your repo's shape.
Conclusion: write down version and harness first
As of October, the coding and knowledge-work center moved to Terminal-Bench 4.0, FrontierCode, CursorBench, AA-Briefcase v1.1, and GDPval-AA v2.1. Opus 5.5 leads broadly, with Astra pushing back in terminal, science, and document reasoning.
But write this first, before any number: version, harness, effort, and run count. A score without those four is not comparable. Make the final call with your own 5 repo questions, not public scores. That is the conclusion, and the cheapest consulting you will get.
Related posts
- SWE-bench vs Terminal-Bench vs ProgramBench: what differs
- Why AI benchmark scores suddenly change
- What is ProgramBench? The rebuild-everything test
References
- Anthropic: Introducing Claude Opus 5.5
- Artificial Analysis: Claude Opus 5.5 takes the top spot
- Artificial Analysis: Benchmarking GPT-6 Astra
- Artificial Analysis: Gemini 4 Argon
- Artificial Analysis: Models Leaderboard
Go deeper with a course
If you want to go beyond reading benchmarks and build harnesses and verification loops tuned to your own project, practice measuring success, cost, and time in a real repo.