Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
Junlin Yang, Che Jiang, Yu Fu, Tianwei Luo, Can Ren, Weizhi Wang, Kaikai Zhao, Hongyi Liu, Yuxin Zuo, Yuru Wang, Yuchen Fan, Kai Tian, Zhenzhao Yuan, Xiaojian Lin, Li Sheng, Rushi Qiang, Guoli Jia, Xingtai Lv, Ermo Hua, Dianqiao Lei, Youbang Sun, Ning Ding, Bowen Zhou, Kaiyan Zhang
cs.CL
Submitted: 2026-07-30
Code: https://github.com/FrontisAI/OpenRSI
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this
Terminology
Abstract
Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI. Code: https://github.com/FrontisAI/OpenRSI
Sources
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
- OpenAI Gym
- AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering
- AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
- Measuring AI R&D Automation
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- Toward Autonomous Long-Horizon Engineering for ML Research
- Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
- SELA: Tree-Search Enhanced LLM Agents for Automated Machine Learning
- AutoMLGen: Navigating Fine-Grained Optimization for Coding Agents
- MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery
- AutoGluon-Tabular: Robust and Accurate AutoML for Structured Data
- MLZero: A Multi-Agent System for End-to-end Machine Learning Automation
- ResearchGym: Evaluating Language Model Agents on Real-World AI Research
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- AIRA_2: Overcoming Bottlenecks in AI Research Agents
- Automated Design of Agentic Systems
- MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
- AIDE: AI-Driven Exploration in the Space of Code
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering