On September 17, MIT and Sakana AI researchers posted a paper. SIFT (Self Improvement via Fast Tree-search, arXiv 2609.19526). Coding agents that edit their own prompts, tools, and code to improve themselves.
The recursive self-improvement post covered the era of "AI fixing itself." SIFT answers the next question. Editing works, but figuring out which edit is good costs far too much.
The bottleneck moved from generating to verifying
The old self-evolution shape looks like this.
Generate candidate edits → run all of them on pricey benchmarks → judge by score
Problem: hundreds of candidates mean thousands of CPU hours plus steep bills
(SWE-bench full eval costs are the classic example)
The paper names the bottleneck precisely. Full benchmark runs give reliable signal at high cost; running a small task subset is cheap but noisy. The signal-per-dollar of both is poor.
SIFT adds a middle signal. Before the expensive real eval, an LLM judge ranks candidates with pairwise comparisons, and only the promising ones get truly tested.
The SIFT shape:
generate candidates → pairwise judge bouts → Bradley-Terry strength scores
→ rank-based parent sampling (tree keeps growing)
→ only top nodes get pricey downstream evals (async, in parallel)
The trick is duels instead of grades. Rather than asking "is this patch good," it asks "which of these two patches is better" many times and converts win-loss records into scores. And the search never stalls. While slow evals run, judge signals grow the next branches. A disaggregated pipeline, in the paper's words: search and evaluation decoupled.
Numbers: 35.1% and the judge's share
Polyglot results (225 tasks).
| Setup | Score |
|---|---|
| o3-mini base agent | 14.2% |
| SIFT without judge (o3-mini) | 29.8% |
| DGM (prior best) | 30.7% |
| SIFT plus gpt-5.4 judge (o3-mini) | 35.1% |
| Qwen3-30B base to SIFT | 20.0% to 32.0% |
One row deserves attention. Judge-free SIFT (29.8%) sits below DGM (30.7%). Adding the judge lifts it to 35.1%. The tree structure is not the hero; the judge's contribution is isolated and visible.
Resources matter too. Under 50 CPU hours and 5 wall-clock hours on o3-mini; under 250 CPU hours and 7 hours on Qwen3. Compare that with thousands of CPU hours in older methods and the entry barrier collapses. Improved harnesses also beat base harnesses on other models (gpt-5-mini, gpt-5.4-mini), so these look like design gains rather than overfitting to one model.
Practitioner translation: split into 5 stages
Nobody needs to rebuild the paper. As VentureBeat noted, SIFT sits on the DGM harness, so take the skeleton as 5 stages.
1. Candidate generation: collect edits (prompts, tools, code)
2. Cheap judge: rank with pairwise bouts (no real evals yet)
3. Real eval: test only the top ranks on real tasks
4. Held-out regression: check for regressions on a separate set
5. Promotion: ship only passers, with records
The five-question repo method covers stages 3 and 4, and the router log covers stage 5. The missing piece is stage 2, a cheap judge. One "which of the two is better" judge prompt starts it.
CodeBridge Mini Lab: attach one judge
1. Pick 2 edits (example: system prompt A vs B)
2. Write 1 judge prompt:
"Compare the two candidates as code and pick exactly one winner.
Reasons in 3 lines or fewer."
3. Run 10 bouts and count wins (Bradley-Terry can wait)
4. Test only the winner on 5 real tasks
5. Call it: if the judge winner matches the measured winner,
trust the judge and scale up
It matches the measuring habit from the Pass@1, cost, and time post. Filter with a cheap judge before full evals. That paragraph is SIFT's practitioner edition.
Conclusion: what deserves verifying
Restating the moved question.
Not "can we make better edits" but "what deserves verifying."
When building an agent improvement pipeline, do not scale generation first. Lay down cheap judge, real eval, separate-set regression, and promotion first. Cheap verification brings more attempts. More attempts bring improvement. The order is not reversible.
Further reading
- Recursive self-improvement: how far has it come
- What is ProgramBench: the rebuild-it-all exam
- Why read Pass@1, cost, and time together in coding agent ranks
References
- arXiv: Self Improvement via Fast Tree-search (2609.19526)
- SIFT project page
- VentureBeat: MIT and Sakana AI framework uses an LLM judge (Oct 2, 2026)
- Darwin Gödel Machine (DGM)
Go deeper with a course
To design generate, judge, eval, and ship loops as a structure, this course builds harness, loop, and graph layers exactly like the 5-stage pipeline here.