Coding AI benchmark tables changed a lot in September.

I already wrote a SWE-bench, Terminal-Bench, and ProgramBench comparison. That older baseline no longer reads today's scores well. Terminal-Bench 4.0, AA-Briefcase v1.1, and GDPval-AA v2.1 are now mainstream.

This post organizes the new scoreboard as of October 1. It also covers what matters more than scores: the same model gets different numbers depending on who measures it.

Terminal-Bench 4.0: finish the job in the terminal

Terminal-Bench does not grade chat answers. It runs commands in a terminal and checks whether the task is finished. Key scores on version 4.0 look like this:

Model Score Conditions
Claude Opus 5.5 66.4% Anthropic internal, xhigh, Claude Code harness
Claude Opus 5.5 59.6% Artificial Analysis independent test
GPT-6 Astra 59% Artificial Analysis test
GPT-6 Astra 57.9% OpenAI report, high effort
Claude Opus 5 52.3% Reproduction test

The two Opus 5.5 numbers, 66.4% and 59.6%, are not a typo. Different harnesses and settings produce different results. Anthropic also reports a standard error of plus or minus 2.6 points.

Lesson 1: every score travels with "who measured, with which harness, how many runs"
Memorizing numbers alone misleads you. Write down the conditions too.

Read these 3 coding benchmarks together

Tests that grade code changes themselves were updated too.

Test What it checks Key scores
FrontierCode v1.1 Is the code change merge-worthy Opus 5.5 54.4%, Astra 53.3%, GPT-5.6 Sol 47.5%
CursorBench 4.0 Ambiguous real-world multi-file work Opus 5.5 57.8%, Fable 5.1 51.8%, Opus 5 46.6%
Terminal-Bench-Science 0.1 Terminal-based science research workflows Astra 64.6%, Opus 5.5 58.7%, Opus 5 29.0%

Notice that Astra beats Opus 5.5 on Terminal-Bench-Science. The overall intelligence leader does not win every test. Each test measures something different, so read the test that resembles your work.

Knowledge-work benchmarks: AA-Briefcase and GDPval

Tests for docs, analysis, and decks have new versions too.

Test What it checks Key scores
AA-Briefcase v1.1 Long order-level knowledge work, analysis plus deck quality Opus 5.5 Elo 1822, +143 over Fable
GDPval-AA v2.1 44 real professional tasks Opus 5.5 Elo 1846, Fable 1735, Opus 5 1708
AutomationBench Cross-app office automation Opus 5.5 40.0% (Zapier test), Astra 69% (AA implementation)
GDP.pdf Expert document reasoning, full pass rate Astra 31%, GPT-5.6 Sol 27%

There is a trap here too. Opus at 40% and Astra at 69% on AutomationBench come from different implementers. The Opus number is Zapier's own eval. The Astra number is an Artificial Analysis implementation. The same test name with a different implementation can behave like a different test.

Lesson 2: the same name with a different implementation is a different test
Do not compare "AutomationBench 69% vs 40%" directly.

Also note that Opus 5.5 is the first Anthropic model to win on both analysis and deck quality in AA-Briefcase, while Gemini 4 Argon topped the rubric pass rate at 65% but scored lower on presentation. In long work, analysis skill and presentation skill can diverge.

A 4-item checklist for reading scores

When you see a new scoreboard, write these four items next to every number:

[ ] Bench version: 4.0, v1.1, or v2.1?
[ ] Harness: Claude Code, Codex, or a reference harness like Stirrup?
[ ] Effort: which of low / medium / high / xhigh / max?
[ ] Run count: single run or 3-5 run average (is there a standard error)?

"Opus 5.5 Terminal-Bench 66.4%" alone is half the story. Add "Anthropic internal, xhigh, Claude Code harness, standard error plus-minus 2.6" so your future self is not confused. This matches the principles in how to read bench scores and why scores change.

CodeBridge Mini Lab: make 5 questions for your own repo

Public benchmarks are someone else's exam. What you really need is 5 questions from your own repo:

1. One recurring bug fix (judged by tests)
2. One refactor across 3+ files (judged by tests plus diff scope)
3. Summarize one doc into one table (a human judges in 5 minutes)
4. Add one small feature with no new dependencies (judged by tests)
5. Explain one failing command (a human judges)

Run each model and effort 3 times, record success rate, time, and cost
→ if the public No. 1 differs from your No. 1, trust yours

This follows the same idea as rebuilding ProgramBench from scratch. You shrink the rebuild-everything test into your repo's shape.

Conclusion: write down version and harness first

As of October, the coding and knowledge-work center moved to Terminal-Bench 4.0, FrontierCode, CursorBench, AA-Briefcase v1.1, and GDPval-AA v2.1. Opus 5.5 leads broadly, with Astra pushing back in terminal, science, and document reasoning.

But write this first, before any number: version, harness, effort, and run count. A score without those four is not comparable. Make the final call with your own 5 repo questions, not public scores. That is the conclusion, and the cheapest consulting you will get.

Related posts

References

Go deeper with a course

If you want to go beyond reading benchmarks and build harnesses and verification loops tuned to your own project, practice measuring success, cost, and time in a real repo.