Leaderboards produce strange moments. The same model shows different scores in an old article and on today's site.
Did the model secretly get weaker? Possible. But check something else first.
Did the test itself change?
Benchmarks get version bumps too
Like software, benchmarks change questions, environments, graders, and time limits.
The Artificial Analysis Coding Agent Index moved from Terminal-Bench 2.1 to 4.0 in v1.5 in September 2026. Terminal-Bench 4.0 uses 66 harder tasks, a revised environment and verifier, and a new compute and time budget.
DeepSWE moved from v1.0 to v1.1 at the same time.
So these similar-looking numbers are risky to compare directly.
Coding Agent Index v1.4: 55
Coding Agent Index v1.5: 51
You cannot call that a 4-point drop. The test bundle changed.
Why do tests keep changing?
Four common reasons drive updates.
- The test got too easy to separate models
- Flaky tests or grading bugs surfaced
- Authors want tasks closer to real work
- Public questions risk contamination or answer lookup
Older benchmarks also invite overfitting. Models and agents tune to familiar test shapes.
CodeBridge Mini Lab: comparing two news articles
When you read model comparisons, jot down these five lines. Bad comparisons drop sharply.
Model:
Benchmark:
Benchmark version:
Harness / settings:
Measured date:
For example:
Model: GPT-X
Benchmark: Terminal-Bench
Version: 2.1
Agent: A
Date: 2026-08
versus:
Model: GPT-X
Benchmark: Terminal-Bench
Version: 4.0
Agent: B
Date: 2026-09
Same model name is not enough. Never compare those two rows head to head.
Same benchmark name can hide different grading
Version changes go beyond question lists.
In Coding Agent Index v1.5, DeepSWE v1.1 grades the committed patch in a separate verifier environment, not the agent's working environment. That matters. It reduces false passes where an agent accidentally breaks its local setup in a way that makes tests pass.
So a benchmark version really means this bundle.
Tasks
+ Environment
+ Tools
+ Time budget
+ Verifier
+ Scoring rule
How to write scores in your own blog
Bad example:
Model A scored 57 on a coding benchmark.
Better:
In the September 2026 Artificial Analysis Coding Agent Index v1.5, Model A with a specific agent setup scored 57.
Dates and versions keep your post meaningful after the next benchmark update.
Conclusion: read the version before the number
AI leaderboards update fast. Posts that survive do not just copy numbers.
They record which test, which version, and which environment produced the score.
When a score moves, check whether the exam changed before you blame the model.
Further reading
- How to read AI benchmark scores correctly
- SWE-bench vs Terminal-Bench vs ProgramBench
- What is AI benchmark reward hacking?
References
Go deeper with a course
If you want to stop chasing headlines and pick tools by task fit, this course covers practical, situation-based AI selection.