When you hear "AI coding benchmark," you usually picture an agent fixing an issue in an existing repository.
But real development includes a completely different kind of work:
What if there is no source code, and you must rebuild the program from its behavior and documentation alone?
ProgramBench evaluates exactly that question.
ProgramBench gives you two things
The setup of ProgramBench is strict:
Input
- compiled binary
- documentation
Goal
- implement a codebase that behaves like the original program
The model cannot find an existing source patch. It must observe the program's interface and behavior, design an architecture, and implement it.
It measures a different skill than SWE-bench
The SWE-bench family usually starts from an existing repository:
Existing codebase
+ GitHub issue
→ locate relevant code
→ patch
→ tests
ProgramBench starts from somewhere else entirely:
Binary + Docs
→ infer behavior
→ design architecture
→ implement
→ reproduce behavior
So you should never mix the two exams into one word like "coding score."
Why from-scratch ability matters
As AI coding agents get used more widely, work beyond simple patches keeps growing:
- reimplementing old internal tools
- building CLI-compatible implementations
- rewriting a prototype as production code
- developing new services from a specification
- cloning the behavior of a reference app
These tasks depend less on repository search and more on requirement inference and system design.
CodeBridge Mini Lab: reimplement a tiny CLI as a black box
If running the full ProgramBench is too much, you can build a tiny version of it.
First, prepare one simple CLI program:
$ original-tool normalize " Hello World "
hello world
$ original-tool count "a,b,c"
3
Give the agent no source code — only this:
- a runnable binary or sample endpoint
- the --help documentation
- a few examples
Then ask it to build a compatible program.
Validate with hidden inputs:
seen examples → easy to pass
unseen edge cases → real understanding of the behavior
This experiment is a good way to separate "memorized and copied the examples" from "inferred the rules."
Long tasks make the harness matter more
From-scratch implementation has many steps:
Explore behavior
→ Write spec
→ Design
→ Implement
→ Test
→ Find mismatch
→ Refine
If the model tries to finish all of this in one long response, it easily misses errors in the middle.
So the execution structure matters: the agent must save progress, analyze test failures, and fix them again.
Conclusion: "AI that fixes code" and "AI that builds software" need different tests
You cannot assume a model that is great at bug fixes will always design a great architecture from an empty project.
ProgramBench is interesting because it looks at AI coding more broadly:
Beyond understanding and fixing existing code, it asks whether AI can understand behavior and rebuild it into a finished program.
Further reading
- SWE-bench vs Terminal-Bench vs ProgramBench
- What is SWE-bench Multimodal v2?
- How many hours of work can an AI agent do alone?
References
Go deeper with a course
If you want to practice the harness structures that carry long coding tasks like this, a guided course is the fastest way to build the habit.