EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making
cs.AI
Submitted: 2025-08-13
Updated: 2026-08-26
Code: https://github.com/opendilab/DI-star
Project page: https://oasismodel.github.io
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Complex decision-making often requires agents to progress through intermediate tasks rather than solve the final target directly.
Terminology
Abstract
Complex decision-making often requires agents to progress through intermediate tasks rather than solve the final target directly. Existing LLM self-refinement methods typically iterate on a fixed target, while curriculum-learning methods often rely on hand-designed schedules, training-time optimization, or domain-specific difficulty metrics. To smooth out the learning curve with adaptive curriculum desgin, we introduce EvoCurr, a general inference-time framework that co-evolves curricula and executable policies. A Designer proposes verifiable intermediate tasks from the latest accepted progress and recent failures, while a Solver generates or trains policies for these tasks. A task-policy pair is accepted only when it passes hard feasibility checks and reaches a specified performance threshold meanwhile an accepted-floor constraint preserves the latest verified checkpoint and prevents failed harder attempts from overwriting mastered behavior. We evaluate EvoCurr across code-as-policy and closed-loop MARL settings on StarCraft II super-late-game micromanagement, hardest stress test, where EvoCurr achieves 96.7--99.2% macro-average win rates across three LLM backbones without LLM parameter training. The performance improvments on other experiments, Flatland, POGEMA, Overcooked, and LiveCodeBench Hard, reflect the generality of our EvoCurr.
Sources
- LLM-PySC2: Starcraft II learning environment for Large Language Models
- An Introduction of mini-AlphaStar
- AVA: Attentive VLM Agent for Mastering StarCraft II
- Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach
- GPT-4 Technical Report
- Automatic Curriculum Learning For Deep RL: A Short Survey
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
- AgentSquare: Automatic LLM Agent Search in Modular Design Space
- AdaPlanner: Adaptive Planning from Feedback with Language Models
- StarCraft II: A New Challenge for Reinforcement Learning
- GenSim: Generating Robotic Simulation Tasks via Large Language Models
- Genie: Generative Interactive Environments
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- GameGen-X: Interactive Open-world Game Video Generation
- Evaluating Large Language Models Trained on Code
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- SMAC-R1: The Emergence of Intelligence in Decision-Making Tasks
- SMAC-Hard: Enabling Mixed Opponent Strategy Script and Self-play on SMAC
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection