Model A scores 96. Model B scores 97. Can you declare B the better model?

If the test is still hard, maybe. But if most top models crowd into the high 90s, the story changes.

That state is called benchmark saturation.

Saturated tests stop separating models

Think of a school exam.

If students spread from 50 to 80, the test separates skill levels. If everyone nears 100, a one-point gap means little.

AI benchmarks behave the same way.

Early days
Model A 45
Model B 62
Model C 78

After saturation
Model A 96
Model B 97
Model C 98

In the second case, run-to-run noise, sampling, and grading error can move rankings more than real skill gaps.

Saturation does not mean "models are perfect"

This is the most common mistake.

99% on a benchmark = 99% on real work

That equation fails.

A benchmark measures one task distribution, not all of reality. Acing an old test does not prove skill at long projects, UI control, huge codebases, or brand-new tools.

That is why new evals move toward longer tasks, agentic workflows, private test sets, and real deliverables.

CodeBridge Mini Lab: check the spread before the rank

When you find a leaderboard, skip first place. Look at the top-10 range first.

Top 10 range: 94 ~ 98

If it is this narrow, ask one more question.

Does this benchmark still separate frontier models?

Compare with:

Top 10 range: 31 ~ 68

That test likely still reveals big model gaps.

Range alone cannot prove saturation. But it is a good warning signal.

Why new benchmarks keep getting harder

The 2026 eval trend favors these shapes over simple Q and A:

  • Knowledge work across hundreds or thousands of files
  • Automation that drives real SaaS tools
  • Multi-step terminal tasks
  • Agent tasks that read and edit whole repos
  • Rebuilds from binaries plus docs alone

The direction is clear. Measure whether the job gets finished, not whether one question gets answered.

Saturated benchmarks are still useful

Old tests still help catch regressions.

If a new model suddenly drops far below its generation on an old benchmark, suspect a real capability loss.

They just carry less signal for fine frontier rankings.

Use old tests as health checks. Use new tests for head-to-head picks.

Conclusion: good benchmarks must keep getting harder

Models improve, so tests must improve too.

Benchmark replacements and new versions feel confusing. They are natural.

Do not stop at "what is the score?" Ask the next question too.

Does this test still show real differences between current models?

That second question turns numbers into decisions.

Further reading

References

Go deeper with a course

If you want a practical way to choose AI tools without getting fooled by crowded rankings, this course teaches situation-based selection.