When you hear "AI coding benchmark," you usually picture an agent fixing an issue in an existing repository.

But real development includes a completely different kind of work:

What if there is no source code, and you must rebuild the program from its behavior and documentation alone?

ProgramBench evaluates exactly that question.

ProgramBench gives you two things

The setup of ProgramBench is strict:

Input
- compiled binary
- documentation

Goal
- implement a codebase that behaves like the original program

The model cannot find an existing source patch. It must observe the program's interface and behavior, design an architecture, and implement it.

It measures a different skill than SWE-bench

The SWE-bench family usually starts from an existing repository:

Existing codebase
+ GitHub issue
→ locate relevant code
→ patch
→ tests

ProgramBench starts from somewhere else entirely:

Binary + Docs
→ infer behavior
→ design architecture
→ implement
→ reproduce behavior

So you should never mix the two exams into one word like "coding score."

Why from-scratch ability matters

As AI coding agents get used more widely, work beyond simple patches keeps growing:

  • reimplementing old internal tools
  • building CLI-compatible implementations
  • rewriting a prototype as production code
  • developing new services from a specification
  • cloning the behavior of a reference app

These tasks depend less on repository search and more on requirement inference and system design.

CodeBridge Mini Lab: reimplement a tiny CLI as a black box

If running the full ProgramBench is too much, you can build a tiny version of it.

First, prepare one simple CLI program:

$ original-tool normalize " Hello  World "
hello world

$ original-tool count "a,b,c"
3

Give the agent no source code — only this:

- a runnable binary or sample endpoint
- the --help documentation
- a few examples

Then ask it to build a compatible program.

Validate with hidden inputs:

seen examples       → easy to pass
unseen edge cases   → real understanding of the behavior

This experiment is a good way to separate "memorized and copied the examples" from "inferred the rules."

Long tasks make the harness matter more

From-scratch implementation has many steps:

Explore behavior
→ Write spec
→ Design
→ Implement
→ Test
→ Find mismatch
→ Refine

If the model tries to finish all of this in one long response, it easily misses errors in the middle.

So the execution structure matters: the agent must save progress, analyze test failures, and fix them again.

Conclusion: "AI that fixes code" and "AI that builds software" need different tests

You cannot assume a model that is great at bug fixes will always design a great architecture from an empty project.

ProgramBench is interesting because it looks at AI coding more broadly:

Beyond understanding and fixing existing code, it asks whether AI can understand behavior and rebuild it into a finished program.

Further reading

References

Go deeper with a course

If you want to practice the harness structures that carry long coding tasks like this, a guided course is the fastest way to build the habit.