Real GitHub issues do not always arrive as friendly text.

"On mobile the button looks like this"
+ screenshot.png

Sometimes you must compare a design mockup against the current screen and fix the gap.

SWE-bench Multimodal exists to measure that reality.

How it differs from classic SWE-bench

Classic SWE-bench uses real GitHub issues and repositories. It asks a model to write a code patch that resolves the problem.

The multimodal version adds visual inputs:

  • bug screenshots
  • UI issue images
  • design mockups and wireframes
  • diagrams describing desired behavior
  • error information shown on screen

So an agent must do more than read text. It must interpret images, then connect them to code.

Why v2 keeps 480 tasks

SWE-bench Multimodal v2, released on September 1, 2026, keeps 480 reproducible tasks. Flaky tests and hard-to-grade cases were removed. The test split and evaluation tooling are open source.

The smaller count is not the point.

It tests visual understanding plus repo navigation plus code editing plus test passing in one workflow.

CodeBridge Mini Lab: fix UI from one screenshot

You do not need to run the full benchmark. Build a small version yourself.

Prepare:

1. A small web project
2. A current UI screenshot
3. A target UI mockup
4. Playwright or existing UI tests

Ask the agent like this:

Compare the current screen with the target mockup.
Explain the differences, then edit only the needed files.
Run the tests and report the results.

Watch for these four signals:

- Did it spot the wrong visual difference?
- Did it touch JS when only CSS needed a fix?
- Did it invent features missing from the screenshot?
- Did it break existing breakpoints?

Those four reveal the quality of a multimodal coding agent surprisingly well.

Good vision does not mean good fixes

You should separate these skills:

Vision
→ recognize what is different

Repository reasoning
→ find where to change

Coding
→ create the patch

Verification
→ confirm the fix worked

An agent can nail the first step and still fail the task. That is why a multimodal benchmark differs from a pure vision benchmark.

Why the agent harness matters for UI work

Screenshot-based tasks depend heavily on the feedback loop:

Screenshot
→ Analyze
→ Edit
→ Run app
→ Capture again
→ Compare

Re-checking the rendered screen beats generating code once and stopping.

So progress in multimodal coding connects to browser and computer tools, and to harness design, not only to model vision.

Conclusion: coding AI is no longer a code-only developer

Real software work mixes logs, terminals, docs, browsers, and images.

SWE-bench Multimodal v2 matters because it evaluates real software issues that include visual information.

When you compare AI coding agents, add one task like this next to code-generation tests:

Can it find the cause from a screenshot, fix it, and verify the fix on screen?

Related posts

References

Go deeper with a course

If you want to build screenshot-to-fix loops with tests and harness feedback in a real repository, learn by doing.