Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning
cs.LG, cs.AI, cs.CL
Submitted: 2026-07-24
Updated: 2026-09-17
Comments: COLM 2026
Code: https://github.com/algorithmicsuperintelligence/openevolve
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains.
Terminology
Abstract
Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains. We hypothesize that the success of such evolution frameworks hinges on meta-skills, such as self-reflection with environment feedback, that enable effective multi-round refinement, yet are largely neglected by traditional post-training. To bridge this gap, we present MetaEvolve, a framework designed to develop these meta-skills via a data synthesis pipeline, evolution-aware reinforcement learning (RL), and inference-time evolutionary search. Concretely, we ground MetaEvolve in coding, where program execution provides natural, continuous reward signals beyond binary correctness. Building on these signals, we synthesize evolution trajectories as training data, each containing a current program, its fitness score (combining correctness and efficiency), and a history of prior attempts, and train the model via RL with verifiable rewards derived from test case execution. By training on large-scale code data, we aim to inspire generalizable domain-agnostic meta-skills that can transfer broadly to open-ended problems where such rich training signals are scarce. Across seven coding benchmarks, MetaEvolve outperforms the strongest baseline by 10.01% absolute on in-distribution tasks and 24.12% on out-of-distribution tasks. On open-ended algorithm optimization problems entirely outside the training domain, it further achieves a 46.9% relative improvement. These results demonstrate that explicitly cultivating self-evolution meta-skills offers a principled path toward more capable and autonomously self-evolving AI.
Sources
- Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
- Constitutional AI: Harmlessness from AI Feedback
- Qwen3-Coder-Next Technical Report
- Process Reinforcement through Implicit Rewards
- Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
- DeltaEvolve: Accelerating Scientific Discovery through Momentum-Driven Evolution
- TACO: Topics in Algorithmic COde generation dataset
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- OpenAI o1 System Card
- AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs?
- CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- Can Language Models Solve Olympiad Programming?
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Large Language Model Reasoning Failures
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks