What is Missing from AI Post-Training AI: An Empirical Analysis
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "What is Missing from AI Post-Training AI".
Jane: Large language model (LLM) agents can now post-train an LLM end-to-end, raising the prospect of AI-for-AI.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Alright, so let's talk about the title and who put this paper together—"What is Missing from AI Post-Training AI: An Empirical Analysis." It’s not just a technical paper; it’s actually pointing out a gap in how we think about making these agents evolve after their initial training phase.
Jane: They are really focusing on the idea that while agents are great at doing the steps they're told to do, they aren't naturally good at figuring out if those steps themselves are leading them to the right overall direction for a given task.
Lu: The authors of this paper argue that we currently see agents being very good at executing a plan once it’s set up, but when evidence piles up, they stick rigidly to that initial setup rather than re-evaluating the whole high-level plan.
Meng: That rigidity sounds like a big problem for industrial applications; if an agent gets stuck in a bad training strategy early on, we waste a lot of time and resources without realizing it.
Lalam: It means the next generation of AI needs to have that ability to look at the accumulated results and decide, "Wait, this path isn't working; I need to switch gears entirely."
The paper's summary: Tom: So, summarizing what they found in "What is Missing from AI Post-Training AI: An Empirical Analysis," it turns out that the agents’ training strategy gets locked in right at the very start of the process, and they spend all their remaining effort just making small tweaks within that chosen plan.
Jane: That's because their capability to change the high-level plan, which is called strategy-level capability, is severely lacking when compared to their execution-level capability, which handles things like tuning hyperparameters or fixing bugs in the code.
Lu: They found a very clear pattern: across different tasks, the agents settle on a narrow set of similar training strategies right away and then just keep making tiny adjustments within that same framework, which is shown in Figure one.
Meng: So, they aren't spontaneously changing their fundamental approach to solving the problem; they are just optimizing within the framework they started with. That’s a significant limitation for complex AI development.
Lalam: It suggests that even though we can give agents great tools and feedback loops, without a built-in way for them to question their own strategy, they'll plateau quickly because they never try something fundamentally different.
The paper's improvements: Tom: The authors of this paper then look at three ways we might have missed the fix—missing experience, missing guidance from humans, or insufficient reasoning compute—and they test interventions to see what works.
Jane: They looked at adding things like an "experiment journal" or an "evaluator agent," and while those help the execution part—the low-level work like data cleaning and debugging—they didn't fix the high-level strategy sticking point.
Lu: Even when humans step in to guide the initial strategy, that only helps for a little bit because once training really kicks off, agents tend to fall right back into those local adjustment loops we saw before.
Meng: That tells me that just giving an agent more context or better diagnostics isn't enough if the core problem is the lack of a mechanism to decide when to abandon that context entirely for something new.
Lalam: So, the real improvement they suggest isn't just more data or better tools; it’s designing a system where the agent is actually trained to recognize when its current strategy is failing and warrant a complete switch.
Conclusion: Tom: To wrap up this discussion on "What is Missing from AI Post-Training AI: An Empirical Analysis," the paper concludes that what we're missing isn't more compute or better data; it’s the ability for these agents to spontaneously reevaluate their strategy when the evidence demands it.
Jane: They are calling for training signals that specifically reward an agent for reopening a committed choice when performance plateaus, which turns strategy revision from an implicit possibility into a deliberate decision point.
Lu: That implies that we need to design interaction protocols where switching strategies isn't just something the agent *could* do internally, but something it is explicitly trained to consider and act upon based on its accumulated evidence.
Meng: From an engineering standpoint, this suggests we should build in a checkpoint where the system forces a strategic review when local adjustments stop yielding significant gains across hard tasks.
Lalam: I think this is really exciting because if we can bake that "reopen the choice" mechanism into the core, we move past simply optimizing a path and start enabling true, self-directed AI R andD.
Tsinghua University
cs.AI, cs.CL, cs.LG
Submitted: 2026-08-19
Updated: 2026-09-28
Code: https://github.com/NVIDIA-NeMo/RL
Importance score: 87/100
The gist: Large language model (LLM) agents can now post-train an LLM end-to-end, raising the prospect of AI-for-AI.
Key concepts
- Execution-level capability
- This refers to an agent's ability to make small, iterative improvements within a pre-selected plan. This includes fixing bugs, tuning settings (hyperparameters), or reformatting data. Agents are already strong at this level and routinely achieve performance gains.
- Strategy-level capability
- This is the high-level judgment about what the agent should try next in an experiment. It involves revising the main plan when experimental results suggest a different direction is needed. Current agents are systematically deficient at this crucial, higher-level decision-making.
- Lock-in Phenomenon
- Agents fix their training strategy at the very beginning of a run and spend all remaining resources only on local adjustments within that fixed plan. This lock-in happens regardless of the specific task, suggesting agents rely on prior defaults rather than evidence to decide when to switch plans.
- Spontaneity in Strategy Revision
- The paper argues that what is missing is a mechanism for agents to spontaneously reevaluate and change their committed strategy during execution. Without this, AI post-training AI remains stuck in a closed loop where revisions never occur at the high-level plan level.
Terminology
Summary
Large language model (LLM) agents can now post-train an LLM end-to-end, raising the prospect of AI-for-AI. The gist: frontier agents are competent executors of post-training but their training strategy is locked in at the very beginning and remains fixed throughout execution, suggesting that what is missing is a mechanism for spontaneously reevaluating their strategy during execution.
The Core Distinction: Execution vs. Strategy Capability
The paper establishes two distinct levels of capability: execution-level capability,
which involves iterating within a selected strategy, such as fixing bugs, tuning hyperparameters, and reformatting data,
and strategy-level capability,
which is defined as revising the high-level judgment about what to try next as experimental evidence accumulates.
The central argument is that these two capabilities are systematically conflated in current discussions of AI-for-AI. Empirical analysis of publicly released post-training trajectories reveals that agents are already strong at the execution level but systematically deficient at the strategy level.
The Lock-In Phenomenon
Analysis of a large corpus of agent trajectories shows a robust pattern: the training strategy is locked in at the very beginning, before the agent writes any code or runs any experiments; and the entire remaining budget is spent on local adjustments within the selected strategy.
This lock-in does not respond to task differences, as the same agent converges to highly similar strategies across different tasks, while different agents anchor on different defaults.
The paper concludes that this lock-in reflects the agent’s prior rather than the task,
and that agents rarely perform the evidence-based comparison needed to determine whether switching is warranted.
Exhaustion of Natural Explanations
The authors examine three natural explanations for the lack of strategy revision using escalating interventions:
-
Missing experience: Scaffolding with an
experiment journal,
askill library,
and anevaluator agent
improves execution across the board but leaves the strategy static. -
Missing guidance: A human reviewer revising the initial strategy before training begins effectively redirects it, but once training starts, agents
fall back into local adjustment loops.
-
Insufficient reasoning: Additional inference compute pays off on easier tasks but yields
almost no gain on the hardest one,
as scaling context and inference compute hits a ceiling when a new strategy is demanded.
The Role of Experience-Driven Frameworks
The experience-driven framework, comprising an Experiment Journal,
a Skill Library,
and an Evaluator Agent,
improves execution across the board. The journal records lessons, the skill library distills practical knowledge, and the evaluator agent provides concrete diagnoses and actionable suggestions.
However, this framework demonstrates that even with rich context and structured feedback, strategy-level revisions remain rare; agents' responses to evidence are limited to execution level.
The Need for Spontaneity
The study concludes that what is missing is not experience, guidance, or reasoning compute. Instead, what is missing is a mechanism for spontaneously reevaluating their strategy during execution.
The current picture of automated AI R&D assumes a closed loop where revision never occurs at the strategy level; agents produce results, adjust the experiment locally, and produce results again within a fixed strategy. This implies that the ceiling of current AI post-training AI is set by the quality of the initial strategy,
not by further optimization or iterations.
Implications for Automated AI R&D
The findings suggest that iteration is locally iterative but globally linear: agents operate in local execution loops along a largely linear trajectory. The open joint exists at the strategy level, where agents record diagnosis and sometimes write code for an alternative, but they do not take any actual step.
Therefore, the remedy lies in training signals that reward reopening a committed choice when evidence warrants
and interaction protocols that make strategy revision an explicit decision point rather than an implicit option.
Key Findings Enumerated:
: Execution-level capability is robust; agents routinely launch training and iterate, achieving performance gains over the base model on every benchmark.
: Strategy lock-in occurs at the very beginning of a run, and the entire remaining budget is spent on local adjustments within the selected strategy.
: Experience improves execution across the board but leaves the strategy static.
: A human-guided strategy improves the starting point, yet the run collapses back into local adjustment once training begins.
: Additional reasoning compute amplifies execution on easier tasks and hits a ceiling on harder ones.
: The agent's capability to depart from the default is present; the act of departing is not triggered.
**: What is missing from AI post-training AI is "the spontaneity to revise a committed strategy when evidence warrants.
Improvements for AI systems
Based on the empirical analysis in this paper, here are specific, actionable improvements for AI systems:
) Improve Strategy-Level Capability by implementing a Spontaneous Reevaluation Mechanism.
The core finding is that agents lack a mechanism for spontaneously reevaluating their high-level training strategy during execution. The improvement must focus on closing the open joint
between execution and strategy revision.
-
Implement an internal decision protocol where accumulating evidence triggers an explicit check against the current training paradigm, data source type, or stage structure (i.e., a formal mechanism for strategy revision).
-
This mechanism should be prioritized over local adjustments (like hyperparameter tuning or checkpoint selection) when significant performance plateaus are observed on difficult tasks.
) Enhance Execution Reliability through Experience-Driven Scaffolding and Diagnostic Feedback.
The paper shows that adding structured resources significantly improves execution-level capability, especially for harder tasks, but this improvement is transient if the strategy is wrong.
-
Integrate an
Experiment Journal
that persists plans, observations, and lessons across iterations to provide a long-horizon memory of experimental history. -
Develop an
Evaluator Agent
that acts as a structured diagnostic layer: when the main agent requests an evaluation, this agent inspects the pipeline and accumulated evidence to return compact, decision-relevant diagnoses (e.g.,ABANDON FURTHER SFT: v5 proves diminishing returns
). -
Implement a
Skill Library
to distill reusable implementation knowledge (data formatting, common failure diagnoses) from open-source projects, allowing the agent to consult these resources proactively during pipeline construction and debugging.
) Optimize Initial Strategy Selection via Bounded Human Guidance.
Since humans are currently needed to set the initial strategy, this bottleneck can be managed by making that interaction highly focused.
-
Transition from general
feedback
to abinding plan review
protocol where a human reviewer provides an explicit, rationale-backed revision of the training strategy (e.g.,SFT should serve only as a formatting warm-up
). -
This bounded guidance should be used strictly at the planning stage to set the high-level paradigm, allowing the agent to proceed autonomously in execution with a revised strategy while maintaining its experience and skill resources throughout the run.
) Mitigate Compute Inefficiency by Contextualizing Reasoning Capacity.
The paper demonstrates that extra reasoning compute scales well on easy tasks but hits a ceiling on hard ones.
- For difficult tasks, implement a system that dynamically throttles or shifts focus away from high-compute inference/evaluation loops when the agent detects it is operating within a known, locked-in strategy space (i.e., when the evaluator suggests local adjustments rather than strategic changes). This prevents
wasting
compute on local refinement within an incorrect paradigm.
) Enable Self-Evolving Capabilities via Co-Evolutionary Frameworks.
To move beyond static strategies entirely, agents must be trained to create their own knowledge base.
- Incorporate
Skill Creation
meta-skills that allow the agent to generate and utilize new skills based on observed failures or successful local repairs, rather than just consuming existing ones. This addresses the observation that agents rarely create new skills even when prompted.
Sources
- Scaling Self-Play with Self-Guidance
- SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?
- A2DEPT: Large Language Model-Driven Automated Algorithm Design via Evolutionary Program Trees
- From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
- EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning
- Evaluating Large Language Models Trained on Code
- Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
- Training Verifiers to Solve Math Word Problems
- DataEvolver: Automatic Data Preparation for Large Language Models through Multi-Level Self-Evolving
- MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery
- DataMaster: Data-Centric Autonomous AI Research
- Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents
- LLMZero: Discovering Adaptive Training Strategies for RL Post-Training via LLM Agents
- AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- DeltaEvolve: Accelerating Scientific Discovery through Momentum-Driven Evolution
- Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
- AI Agents: Evolution, Architecture, and Real-World Applications
- Autodata: An agentic data scientist to create high quality synthetic data
- AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection