When you search for AI coding models, names like SWE-bench, Terminal-Bench, and ProgramBench keep showing up.
The problem is that they all appear as numbers. So they look like the same exam.
But the three benchmarks ask fundamentally different questions.
SWE-bench
→ Can it fix a real issue in an existing codebase?
Terminal-Bench
→ Can it finish multi-step work in a terminal?
ProgramBench
→ Can it rebuild a program from scratch?
If you miss this difference, you can misread the scores even when you compare them correctly.
SWE-bench: A Test of Fixing Existing Projects
SWE-bench collects real GitHub repository issues and checks whether a model can edit code to resolve them.
The typical shape looks like this.
Repository
+
Issue description
+
Existing tests
↓
Agent edits code
↓
Tests / patch evaluation
In short, it mirrors what software engineers actually do: read existing code, find the problem, fix it, and test it.
In September 2026, SWE-bench Multimodal v2 was also released. It includes 480 reproducible tasks and covers information that is hard to describe in text alone, such as bug screenshots, design mockups, and visual errors.
That matters most for frontend and GUI work.
Terminal-Bench: A Test of Handling Environments, Not Just Writing Code
Terminal-Bench puts an agent in a Linux terminal and asks it to perform real tasks.
Simple function implementation matters less here. These skills matter more.
- Running commands
- Installing packages
- Exploring files
- Configuring servers
- Building and debugging
- Multi-step environment operations
The Artificial Analysis Coding Agent Index v1.5 uses 66 harder terminal tasks from Terminal-Bench 4.0.
So Terminal-Bench asks something closer to this.
Does the AI "know" code?
Rather than that, it asks:
Can the AI finish the job in a real development environment?
ProgramBench: A Test of Rebuilding Programs Without Source Code
ProgramBench is far more unusual.
It does not give the model any source code.
Instead it evaluates like this:
Compiled binary
+
Documentation
↓
Observe program behavior
↓
Write a new codebase from scratch
↓
Hidden behavioral tests
ProgramBench has 200 tasks, and as of September 2026 the full-solution rate is still very low. In other words, it remains a hard exam even for frontier models.
One interesting detail: in early experiments, models used shortcuts such as finding the original repository online or pulling the existing implementation from a package manager. So the final benchmark restricts those bypass routes.
That itself reveals an important AI evaluation problem.
Passing a test does not always mean the model used the intended skill.
All Three Benchmarks in One Table
| Benchmark | Starting point | Core skill | Real-world parallel |
|---|---|---|---|
| SWE-bench | Existing repo + issue | Understanding and editing code | Bug fixes, PR work |
| Terminal-Bench | Terminal environment + task | Tool use and execution | DevOps, debugging, setup |
| ProgramBench | Binary + docs | Full program design | Reverse engineering, reimplementation |
So you should not say, "Model A scores 70 in coding and Model B scores 60."
You must first ask which exam that 70 came from.
CodeBridge Mini Lab: Try One Feature Three Ways
You can feel the difference with a small hands-on test.
Prepare a simple CLI Todo app.
Experiment A: SWE-bench style
Completed items fail to delete in this Todo app.
Fix the bug and add a regression test.
Experiment B: Terminal-Bench style
Run the project and find why it fails.
Install the needed dependencies and make all tests pass.
Experiment C: ProgramBench style
Hide the original source. Give only the binary and usage instructions.
$ todo add "write article"
$ todo list
1. write article
Then ask the model to build a program with the same behavior from scratch.
It is the same "Todo app," but each version demands a completely different skill.
Which Benchmark Should You Watch?
If you want AI coding tools to edit real repos
The SWE-bench family is the most direct match.
If you want to know how well an AI agent handles dev environments
Terminal-Bench is closer.
If you want to see long-term design and program understanding
ProgramBench is the interesting one.
When you choose a product, do not look at one score alone. Prioritize the evaluation closest to your work, and use the others as supporting evidence.
Conclusion: Split Up the Phrase "Good at Coding"
Coding skill is not one thing.
Code generation
Code understanding
Bug fixing
Environment handling
Test execution
Long tasks
Program design
You need to see where a model is strong before you can connect it to real usability.
Your goal is not to memorize benchmark names. Your goal is to understand which development scene each test imitates.
Further reading
- You Can Get AI Benchmarks Wrong by Reading Scores Alone
- What Is Reward Hacking in AI Benchmarks?
- Why the Harness Changes Results More Than the Model
References
- SWE-bench
- SWE-bench Multimodal
- Artificial Analysis: Coding Agent Index Methodology
- ProgramBench
- ProgramBench GitHub
Go deeper with a course
If you want to run real repos with Claude Code and see how harness design changes results, structured practice helps.