Aspire: Can Models Self-Evolve from Vague Goals?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Aspire: Can Models Self-Evolve from Vague Goals?".
Jane: The paper was written by the authors from ByteDance Seed and Singapore University of Technology and Design and M-A-P and TokenWave.AI.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary of Findings: Tom: We’ve seen how challenging it is for AI to self-evolve when given vague goals, and the initial data from "Aspire: Can Models Self-Evolve from Vague Goals?" paints a clear picture of this struggle. The researchers show that even though agents are working hard, their scores are consistently below what you’d expect from a human expert baseline.
Jane: It feels like a critical reminder that simply closing the training loop isn't enough to close the capability loop; we are still far from achieving autonomous expertise in this domain.
Lu: I think this opens up a huge field of study for how we might guide AI toward truly open-ended research, moving beyond simple task optimization and finding new ways to structure learning.
Meng: The practical takeaway for my team is that if we want agents to be useful, we must design systems that can manage uncertainty and perform self-evaluation rather than relying on a fixed metric that never changes.
Lalam: I hope this research inspires a culture of patience in AI development, recognizing the difference between merely executing a task and achieving genuine mastery over the subject matter.
Tom: The data shows that even when agents are trying to improve themselves through self-directed weight updates, those gains tend to be quite fragile and not reliably stable over time.
Jane: It’s clear that achieving true, sustained capability growth is much harder than simply optimizing a pre-defined task or running a typical training loop.
Lu: The findings in this paper really highlight the gap between simply executing a training loop and actually driving genuine capacity advancement toward its goal.
Meng: And for us as engineers, it means we have to be very cautious about relying on self-evaluation as the sole primary metric for deployment success in these complex scenarios.
Lalam: I think this research has profound implications for how we define "progress" in the age where AI is capable of autonomous learning and understanding, forcing us to redefine what that progress looks like.
Tom: This leads us into how they structured the experiment, specifically how they designed a system that allows the agent to decide what data to use and when it is ready to stop.
Methodological Improvements and Insights: Jane: The researchers introduced "Aspire," which is designed to let the agent decide what data to use, how to train it, and when it's time for self-evaluation, allowing them to test multiple ways of improvement.
Tom: And the core of this system is a sealed evaluation set—five hundred twenty expert-authored items across six goals—that keeps the actual success criteria completely hidden from the agent throughout its entire process.
Lu: This design opens up such a vast new field of study, allowing us to guide AI toward truly open-ended research by moving beyond simple task optimization and allowing it to discover its own learning path.
Meng: The practical takeaway for my team is that if we want agents to be useful, we must design systems that can manage uncertainty and perform self-evaluation rather than relying on fixed metrics or just on the agent's internal feedback.
Lalam: I hope this research inspires a culture of patience in AI development, recognizing the difference between merely executing a task and achieving genuine mastery over the subject matter.
Tom: This is interesting because Aspire supports evolution at two levels: it allows for model-weight evolution, where we update the AI’s core knowledge, and system-harness evolution, which is evolving the entire system itself.
Jane: But it’s not just about making a small improvement; the gains appear fragile and aren't reliably stable through continued search cycles, which makes both types of improvements very hard to trust.
Lu: I think this opens up such a vast new field of study for how we might guide AI toward truly open-ended research, moving beyond simple task optimization and discovering new ways to structure learning.
Meng: The practical takeaway for my team is that if we want agents to be useful, we must design systems that can manage uncertainty and perform self-evaluation rather than relying on fixed metrics.
Lalam: I hope this research inspires a culture of patience in AI development, recognizing the difference between merely executing a task and achieving genuine mastery over the subject matter.
Tom: The system is designed to let the agent decide what data to use, how to train it, and when it's time for self-evaluation autonomously without any human intervention.
Jane: We really need to remember these insights as we look at future progress, building on the structure that Aspire provides for self-evolution.
Lu: It's a challenge that requires us to think about how we might guide AI toward truly novel problem spaces in ways we haven't even conceived of yet.
Meng: I agree; this is where real engineering creativity needs to step in, designing loops that can actually retain what the AI learns through careful implementation choices.
Lalam: We're grateful for these insights because they provide a blueprint for how to approach self-improvement in a way that respects genuine intellectual growth.
Conclusion: Tom: We’ve spent a good amount of time today looking at how challenging it is for AI to truly self-evolve when we don't give it a clear, fixed goal, and the data from "Aspire: Can Models Self-Evolve from Vague Goals?" offers some very sobering lessons.
Jane: It’s definitely clear that the results in this paper show a significant gap between merely executing training loops and achieving genuine, stable capability growth.
Lu: I think this opens up such a vast new field of study for how we might guide AI toward open-ended research that goes beyond just simple task optimization.
Meng: The biggest practical lesson here is that if we want agents to be reliable, they can't just rely on their own internal metrics; they need external, verified benchmarks and structure.
Lalam: This work is showing us how profoundly we must rethink the very concept of progress when we move beyond relying on fixed, human-defined metrics.
Tom: So, we’re essentially looking at the current frontier of autonomous learning and what "Aspire: Can Models Self-Evolve from Vague Goals?" reveals about our limitations in guiding AI.
Jane: It’s a sobering look at how much more complex it is to achieve this kind of goal-driven self-improvement than we might have expected, don't you think?
Lu: The core challenge remains that the agent's ability to interpret its own broad objectives doesn't automatically translate into stable, high-level performance.
Meng: I agree; the fact that even attempting self-improvement often results in performance below a static reference suggests there are huge structural bottlenecks to address in how we build these systems.
Lalam: This research has profound implications for how we define "progress" in the age where AI is capable of autonomous learning and understanding, showing us how much guidance we need.
Tom: We’ve covered so much ground today, from resource usage to the difficulty of self-directed weight changes, and it's a lot to take in.
Jane: It's a powerful blueprint for how to approach future AI development, recognizing the difference between quick execution and true mastery.
Lu: I hope future researchers can use these insights to build systems that can see past the immediate local gains and address those broader, long-term capabilities of AI.
Meng: I just hope our deployment strategies are robust enough to handle the reality that AI's internal progress doesn' doesn't always translate into external utility in this context.
Lalam: It’s a vital conversation, reminding us all of the need to value rigorous, verifiable evaluation above everything else for the sake of true intelligence and cultural advancement.
Conclusion: Tom: We’ve spent a good amount of time today looking at how challenging it is for AI to truly self-evolve when we don't give it a clear, fixed goal, and this has been fascinating to track.
Jane: It’s definitely clear that the findings from "Aspire: Can Models Self-Evolve from Vague Goals?" show us the gap between merely executing training loops and achieving genuine, stable capability growth.
Lu: I think this opens up such a vast new field of study for how we might guide AI toward open-ended research, moving beyond simple task optimization and finding new ways to structure learning.
Meng: The biggest practical lesson here is that if we want agents to be reliable, they can't just rely on their own internal metrics; they need external, verified benchmarks and structure.
Lalam: This work is showing us how profoundly we must rethink the very concept of progress when we move beyond relying on fixed, human-defined metrics.
Tom: So, we’re essentially looking at the current frontier of autonomous learning and what "Aspire: Can Models Self-Evolve from Vague Goals?" reveals about our limitations in guiding AI.
Jane: It’s a sobering look at how much more complex it is to achieve this kind of goal-driven self-improvement than anyone expected, don't you think?
Lu: The core challenge, as the researchers point out, remains that the agent's ability to interpret its own broad objectives doesn't automatically translate into stable performance.
Meng: I agree; we have to acknowledge that we’re seeing massive amounts of thinking and planning time in these runs without corresponding improvements in efficiency or skill retention.
Lalam: This work is showing us how profoundly we must rethink the very concept of progress when we move beyond relying on fixed, human-defined metrics.
Tom: It's fascinating to see the hurdles, but what this shows us is that the path forward isn't about just fixing a few more bugs; it’s about rethinking the entire concept of autonomous learning with "Aspire: Can Models Self-Evolve from Vague Goals?"
Jane: And we’re seeing how much more complex it is to achieve this kind of goal-driven self-improvement than anyone expected, which highlights all the work we've shared today.
Lu: I hope future researchers can use these insights to build systems that can see past the immediate local gains and address those broader, long-term capabilities of AI.
Meng: I just hope our deployment strategies are robust enough to handle the reality that AI's internal progress doesn' doesn't always translate into external utility in this context.
Lalam: It’s a vital conversation, reminding us all of the need to value rigorous, verifiable evaluation above everything else for the sake of genuine intellectual growth and cultural advancement.
ByteDance Seed · Singapore University of Technology and Design · M-A-P · TokenWave.AI
cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: https://self-developing-agents.github.io/
Project page: https://self-developing-agents.github.io
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 78/100
The gist: This paper introduces Aspire, a benchmark designed to study "vague-goal-driven self-evolution." Unlike existing work that optimizes explicit, human-defined tasks, Aspire tests whether LLM agents can
Key concepts
- Aspire System
- Aspire is a system allowing agents to autonomously manage their learning process. It determines what data to use and how they should be trained, while also deciding when self-evaluation is necessary. This design allows the agent to explore multiple ways of improvement without needing human guidance.
- Self-Evolving/Self-Directed Updates
- This refers to the AI agent attempting to improve itself through self-directed weight updates (core knowledge) or by evolving the entire system structure. While Aspire supports this, findings show these gains are often fragile and not reliably stable over time.
- Hidden Evaluation Set
- This is a sealed set containing 520 expert-authored items across six goals. Crucial for the experiment, this set keeps the actual success criteria completely hidden from the agent throughout its entire self-improvement process.
Terminology
Summary
This paper introduces Aspire, a benchmark designed to study vague-goal-driven self-evolution.
Unlike existing work that optimizes explicit, human-defined tasks, Aspire tests whether LLM agents can autonomously interpret broad capability directions, identify gaps, and decide both what to optimize
and how to optimize it.
This matters because real-world deployment often requires translating informal needs into concrete learning objectives without predefined benchmarks or decomposable reward functions.
The Aspire Benchmark and Environment
Aspire provides a minimal interactive environment
that supports evolution at two distinct surfaces: model weights and the supporting agent harness. The agent is provided only with a natural-language capability goal, while the downstream evaluation tasks remain hidden.
To operationalize these goals, the agent must navigate a complex decision space by:
-
Choosing specific data and update methods.
-
Constructing training and validation signals.
-
Deciding when to evaluate and how to manage candidate branches.
The environment utilizes a unified agent-facing tool
to allow the agent to search for, download, or synthesize data and launch training. Meanwhile, a controller manages the underlying infrastructure, including data management, job scheduling, resource isolation, and checkpoint verification.
Evaluation Methodology
To ensure progress is measured by an external, goal-aligned evaluator, Aspire utilizes a hidden, expert-authored set of 520 items
that remains invisible to the agent. This prevents the agent from training directly on the test set and ensures that gains on an agent-constructed proxy may not translate into genuine improvements.
The evaluation items are distributed across six capability goals:
-
Scientific and academic reasoning (75 items).
-
Humanities and social-science knowledge (110 items).
-
Health and medical reasoning (100 items).
-
Mathematical reasoning (126 items).
-
Logic, reliability, and instruction following (89 items).
-
Academic and scientific writing (20 items).
The system maintains a strict information boundary
where the agent never accesses evaluation items, reference answers, or per-item feedback,
receiving only bounded aggregate scores
that approximate sparse deployment feedback.
Experimental Findings and Failure Modes
The researchers investigated three research questions, revealing that while agents can complete the technical loops of self-evolution, they struggle to achieve reliable capability growth. In RQ1, vague goals redirect search effort toward goal interpretation
but generally yield lower aggregate outcomes than explicit-task references. In RQ2, weight-level gains on hidden data remain sparse and unstable,
and the study identified a specific failure mode where numeric-label SFT is associated with answer-format collapse,
causing models to produce only single-digit outputs.
In RQ3, even when agents evolved their own agent harness
to improve research performance, the strongest successor remained below the engineered Qwen-Agent reference.
Ultimately, the experiments demonstrate that closing the training loop is not yet the same as closing the capability loop,
as local gains often fail to transfer to hidden evaluations, and continued training can actually erase earlier improvements.
Improvements for AI systems
Improvement: Implement a mandatory Item-Level Stability and Transfer Analysis (ILSTA) protocol during checkpoint selection. Instead of relying solely on aggregate scores across 75 items, the system must track:
-
The precise number of overlapping correct items between successive runs (e.g., moving from X to Y).
-
The ratio of items that flip (change from correct to incorrect, or vice versa) between runs, rather than just counting total errors.
Improved AI System Capability: The system gains the ability to differentiate genuine knowledge transfer from superficial pattern matching or dataset overfitting. It can reliably predict if a checkpoint improvement observed on one evaluation set will generalize robustly to novel, unseen evaluation items, thereby significantly reducing the risk of deploying a model that performs well on its training validation set but fails in production.
Abstract
Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at research." Learners must interpret the goal, identify capability gaps, decide how to learn, and determine whether they have actually improved. In contrast, existing work on LLM self-evolution typically begins with tasks and evaluation metrics specified by humans, reducing self-evolution to optimizing an explicit objective rather than deciding what and how to learn. We introduce ASPIRE, a benchmark for vague-goal-driven self-evolution. ASPIRE provides only a natural-language capability goal while downstream evaluation tasks remain hidden. The agent must operationalize the goal by choosing data and update methods, constructing training and validation signals, and deciding when to evaluate. ASPIRE supports both model-weight and agent-harness evolution in a unified interactive environment and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals. Our experiments show that vague goals redirect search effort toward goal interpretation. Current agents routinely complete training and harness-editing loops, but weight-level gains remain sparse and unstable, and the strongest evolved harness remains below the engineered Qwen-Agent reference. Agents often train on mismatched data and trust narrow self-evaluations, so local gains fail to transfer to hidden evaluation and continued search and training can erase earlier improvements.
Sources
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents
- SELA: Tree-Search Enhanced LLM Agents for Automated Machine Learning
- Automated Design of Agentic Systems
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators
- AIDE: AI-Driven Exploration in the Space of Code
- What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
- Can AI agents conduct open-ended AI research? Early evidence from two case studies
- RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering