Self Improvement via Fast Tree-search
cs.AI, cs.LG
Submitted: 2026-09-17
Updated: 2026-09-17
License: http://creativecommons.org/licenses/by/4.0/
The gist: Coding agents can recursively modify their own implementations, forming a loop of self-improvement.
Terminology
Abstract
Coding agents can recursively modify their own implementations, forming a loop of self-improvement. While prior work shows this can boost performance on coding benchmarks, existing approaches are costly and compute-intensive. We introduce a simple, sample-efficient self-improvement framework that significantly improves coding performance under strict budget constraints. We identify evaluation of candidate self-modifications as the main runtime bottleneck since prior approaches estimate their effectiveness by re-running a subset of benchmark tasks with the modified agent, which is time-consuming. We introduce Recursive Self Improvement via Fast Tree-search (SIFT), which augments these downstream task evaluations with an LLM-as-a-judge signal that performs pairwise comparisons between candidate patches, where the win-loss record is aggregated with a regularized Bradley-Terry model, and the resulting strength scores drive rank-based parent sampling inside a lightweight disaggregated tree search. Expensive downstream task evaluations are reserved only for the most promising nodes. Using a fully disaggregated tree search pipeline, the judge scores provide intermediate signal to guide exploration on promising candidate patches without being bottlenecked by slow evaluation runs. SIFT outperforms existing tree-search based self-evolution frameworks on the full Polyglot benchmark with significantly lower resource requirements in terms of CPU hours, wall clock time, and API cost.
Sources
- DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
- Measuring Massive Multitask Language Understanding
- Automated Design of Agentic Systems
- The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Goedel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements
- Qwen3 Technical Report
- Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
- Multi-agent Architecture Search via Agentic Supernet
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection