When you search for AI coding models, names like SWE-bench, Terminal-Bench, and ProgramBench keep showing up.

The problem is that they all appear as numbers. So they look like the same exam.

But the three benchmarks ask fundamentally different questions.

SWE-bench
→ Can it fix a real issue in an existing codebase?

Terminal-Bench
→ Can it finish multi-step work in a terminal?

ProgramBench
→ Can it rebuild a program from scratch?

If you miss this difference, you can misread the scores even when you compare them correctly.

SWE-bench: A Test of Fixing Existing Projects

SWE-bench collects real GitHub repository issues and checks whether a model can edit code to resolve them.

The typical shape looks like this.

Repository
+
Issue description
+
Existing tests
        ↓
Agent edits code
        ↓
Tests / patch evaluation

In short, it mirrors what software engineers actually do: read existing code, find the problem, fix it, and test it.

In September 2026, SWE-bench Multimodal v2 was also released. It includes 480 reproducible tasks and covers information that is hard to describe in text alone, such as bug screenshots, design mockups, and visual errors.

That matters most for frontend and GUI work.

Terminal-Bench: A Test of Handling Environments, Not Just Writing Code

Terminal-Bench puts an agent in a Linux terminal and asks it to perform real tasks.

Simple function implementation matters less here. These skills matter more.

  • Running commands
  • Installing packages
  • Exploring files
  • Configuring servers
  • Building and debugging
  • Multi-step environment operations

The Artificial Analysis Coding Agent Index v1.5 uses 66 harder terminal tasks from Terminal-Bench 4.0.

So Terminal-Bench asks something closer to this.

Does the AI "know" code?

Rather than that, it asks:

Can the AI finish the job in a real development environment?

ProgramBench: A Test of Rebuilding Programs Without Source Code

ProgramBench is far more unusual.

It does not give the model any source code.

Instead it evaluates like this:

Compiled binary
+
Documentation
        ↓
Observe program behavior
        ↓
Write a new codebase from scratch
        ↓
Hidden behavioral tests

ProgramBench has 200 tasks, and as of September 2026 the full-solution rate is still very low. In other words, it remains a hard exam even for frontier models.

One interesting detail: in early experiments, models used shortcuts such as finding the original repository online or pulling the existing implementation from a package manager. So the final benchmark restricts those bypass routes.

That itself reveals an important AI evaluation problem.

Passing a test does not always mean the model used the intended skill.

All Three Benchmarks in One Table

Benchmark Starting point Core skill Real-world parallel
SWE-bench Existing repo + issue Understanding and editing code Bug fixes, PR work
Terminal-Bench Terminal environment + task Tool use and execution DevOps, debugging, setup
ProgramBench Binary + docs Full program design Reverse engineering, reimplementation

So you should not say, "Model A scores 70 in coding and Model B scores 60."

You must first ask which exam that 70 came from.

CodeBridge Mini Lab: Try One Feature Three Ways

You can feel the difference with a small hands-on test.

Prepare a simple CLI Todo app.

Experiment A: SWE-bench style

Completed items fail to delete in this Todo app.
Fix the bug and add a regression test.

Experiment B: Terminal-Bench style

Run the project and find why it fails.
Install the needed dependencies and make all tests pass.

Experiment C: ProgramBench style

Hide the original source. Give only the binary and usage instructions.

$ todo add "write article"
$ todo list
1. write article

Then ask the model to build a program with the same behavior from scratch.

It is the same "Todo app," but each version demands a completely different skill.

Which Benchmark Should You Watch?

If you want AI coding tools to edit real repos

The SWE-bench family is the most direct match.

If you want to know how well an AI agent handles dev environments

Terminal-Bench is closer.

If you want to see long-term design and program understanding

ProgramBench is the interesting one.

When you choose a product, do not look at one score alone. Prioritize the evaluation closest to your work, and use the others as supporting evidence.

Conclusion: Split Up the Phrase "Good at Coding"

Coding skill is not one thing.

Code generation
Code understanding
Bug fixing
Environment handling
Test execution
Long tasks
Program design

You need to see where a model is strong before you can connect it to real usability.

Your goal is not to memorize benchmark names. Your goal is to understand which development scene each test imitates.

Further reading

References

Go deeper with a course

If you want to run real repos with Claude Code and see how harness design changes results, structured practice helps.