Real GitHub issues do not always arrive as friendly text.
"On mobile the button looks like this"
+ screenshot.png
Sometimes you must compare a design mockup against the current screen and fix the gap.
SWE-bench Multimodal exists to measure that reality.
How it differs from classic SWE-bench
Classic SWE-bench uses real GitHub issues and repositories. It asks a model to write a code patch that resolves the problem.
The multimodal version adds visual inputs:
- bug screenshots
- UI issue images
- design mockups and wireframes
- diagrams describing desired behavior
- error information shown on screen
So an agent must do more than read text. It must interpret images, then connect them to code.
Why v2 keeps 480 tasks
SWE-bench Multimodal v2, released on September 1, 2026, keeps 480 reproducible tasks. Flaky tests and hard-to-grade cases were removed. The test split and evaluation tooling are open source.
The smaller count is not the point.
It tests visual understanding plus repo navigation plus code editing plus test passing in one workflow.
CodeBridge Mini Lab: fix UI from one screenshot
You do not need to run the full benchmark. Build a small version yourself.
Prepare:
1. A small web project
2. A current UI screenshot
3. A target UI mockup
4. Playwright or existing UI tests
Ask the agent like this:
Compare the current screen with the target mockup.
Explain the differences, then edit only the needed files.
Run the tests and report the results.
Watch for these four signals:
- Did it spot the wrong visual difference?
- Did it touch JS when only CSS needed a fix?
- Did it invent features missing from the screenshot?
- Did it break existing breakpoints?
Those four reveal the quality of a multimodal coding agent surprisingly well.
Good vision does not mean good fixes
You should separate these skills:
Vision
→ recognize what is different
Repository reasoning
→ find where to change
Coding
→ create the patch
Verification
→ confirm the fix worked
An agent can nail the first step and still fail the task. That is why a multimodal benchmark differs from a pure vision benchmark.
Why the agent harness matters for UI work
Screenshot-based tasks depend heavily on the feedback loop:
Screenshot
→ Analyze
→ Edit
→ Run app
→ Capture again
→ Compare
Re-checking the rendered screen beats generating code once and stopping.
So progress in multimodal coding connects to browser and computer tools, and to harness design, not only to model vision.
Conclusion: coding AI is no longer a code-only developer
Real software work mixes logs, terminals, docs, browsers, and images.
SWE-bench Multimodal v2 matters because it evaluates real software issues that include visual information.
When you compare AI coding agents, add one task like this next to code-generation tests:
Can it find the cause from a screenshot, fix it, and verify the fix on screen?
Related posts
- SWE-bench vs Terminal-Bench vs ProgramBench
- In AI coding, the harness can matter more than the model
- What is ProgramBench? AI rebuilds programs from scratch
References
Go deeper with a course
If you want to build screenshot-to-fix loops with tests and harness feedback in a real repository, learn by doing.