AIBuildAI-2.5: Efficient Autonomous AI Model Development Through LLM-Guided Tree Search
cs.CL, cs.AI
Submitted: 2026-09-06
Updated: 2026-09-06
Code: https://github.com/aibuildai/AI-Build-AI
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Autonomous agents that automatically build artificial intelligence (AI) models could broaden access to AI across science and engineering.
Terminology
Abstract
Autonomous agents that automatically build artificial intelligence (AI) models could broaden access to AI across science and engineering. A popular line of such agents frames model building as a code search problem and solves it by tree search, in which each node is a candidate program and the tree grows by generating a child program from a parent, and these agents now approach the capability of experienced AI engineers on realistic benchmarks. However, these agents have three weaknesses in efficiency that have not been fully addressed. First, only a small number of candidates can be executed within a realistic budget, so search rules that rank nodes by executed rewards, such as Monte Carlo-style tree search, rely on few and noisy scores and select the next node to explore less effectively. Second, no resource-aware strategy is used to schedule training jobs, which can lower hardware utilization and training efficiency. Third, every agent call is served by a single powerful model, which inflates inference cost. Here we introduce AIBuildAI-2.5, an agentic system that carries out the tree search with LLM agents and addresses each of the three issues. AIBuildAI-2.5 proposes a novel LLM-guided tree search, in which a judge scores each candidate on its expected improvement, grounding, and feasibility, and a selector ranks the pool of candidates from these scores and the state of the search. In addition, AIBuildAI-2.5 comprises a scheduler that launches training jobs with the current hardware resource status taken into account and a router that assigns lower-cost LLMs to less demanding tasks while reserving the most capable LLM for the most challenging sub-tasks in the AI model building workflow. AIBuildAI-2.5 ranks first on MLE-Bench with a medal rate of 73.3%, and outperforms a strong baseline on six autonomous AI research tasks from AIRS-Bench.
Sources
- AIDE: AI-Driven Exploration in the Space of Code
- R&D-Agent: An LLM-Agent Framework Towards Autonomous Data Science
- AIBuildAI: An AI Agent for Automatically Building AI Models
- Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering
- MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery
- MARS: Modular Agent with Reflective Search for Automated AI Research
- ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning
- InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification
- AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
- AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models
- The FM Agent
- KAPSO: A Knowledge-grounded framework for Autonomous Program Synthesis and Optimization
- Monash Time Series Forecasting Archive
- Character-level Convolutional Networks for Text Classification
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering