Agentic Critical Training
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Agentic Critical Training".
Tom: The gist The proposed Agentic Critical Training (ACT) is a reinforcement learning paradigm that trains large language models to autonomously develop reasoning about action quality by rewarding correct action selection,…
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So we've talked a bit about the Agentic Critical Training paper now and what this means for how we think about training these large language models.
Jane: The authors are basically saying that instead of just teaching an AI to mimic a pre-written reflection, they should train it using reinforcement learning so it learns to autonomously judge which action is better by contrasting choices
Reference <ref:2603.08706#pg1>: .
Lu: It’s about creating genuine self-reflection where the model develops its own reasoning about action quality instead of just imitating text that someone else wrote
Reference <ref:2603.08706#pg2>: .
Meng: This capability to evaluate and compare actions seems like a general mechanism that could be useful for improving decision-making in many different AI systems, even if they aren't strictly agent environments
Reference <ref:2603.08706#pg4>: .
Tom: The paper’s title, Agentic Critical Training, really highlights this shift—it’s about training the model to be critical about its actions in a way that leads to better outcomes
Reference <ref:2603.08706#pg1>: .
Jane: It suggests that by training agents to actively evaluate alternatives through RL, we can get models that are not just following instructions but are actually developing more reflective and capable reasoning abilities
Reference <ref:2603.08706#pg2>: .
Lalam: This move from imitation to genuine self-reasoning is what I’m most excited about for how AI culture develops, because it moves the AI from being a passive responder to an active decision-maker
Reference <ref:2603.08706#pg12>: .
Lu: It’s a path toward models that can handle out-of-distribution situations better and show stronger performance on general reasoning tasks without needing specialized training data
Reference <ref:2603.08706#pg4>: .
Tom: So the big picture is that this approach, Agentic Critical Training, offers a promising direction for making LLM agents more truly reflective and capable in their decision-making processes
Reference <ref:2603.08706#pg2>: .
Conclusion: Tom: So we've been looking at Agentic Critical Training, and now we need to wrap up what this whole thing actually means for the world.
Jane: It’s really about moving past just imitation learning where the AI is just copying what it sees.
Lu: Right, it’s about training these models to actually think critically about which action is better when they have a choice between options.
Meng: So instead of just following a script, the AI learns to pick the superior path based on some internal judgment of quality.
Lalam: It means we're aiming for agents that can do genuine self-reflection, not just parrot back pre-written thoughts.
Tom: Exactly, so the title Agentic Critical Training points to this shift toward self-directed reasoning.
Jane: The authors are showing how they use reinforcement learning to reward the model when it picks the better action among two alternatives.
Lu: They’re pairing an expert action with a generated alternative, and only rewarding the model for picking that expert one.
Meng: So the numbers show it’s getting a solid gain over both imitation and standard reinforcement learning approaches on agent benchmarks.
Lalam: And they also found this method helps general reasoning, even on stuff it wasn't specifically trained for before.
Tom: That’s the big implication—this isn't just about making agents better at specific tasks; it’s about building a more capable foundation for general intelligence.
Jane: It suggests that training models to evaluate action quality directly might be a better way than just supervising them to imitate reflection behaviors.
Lu: It opens up a whole new avenue where the AI develops its own internal logic for what makes an action 'good'.
Meng: I’m curious how this translates practically, though; does it really solve the problem of getting agents to recover when they hit a dead end?
Tom: That’s our next big question—can this self-correction capability actually handle messy, real-world problems without getting stuck in loops?
University of Maryland
cs.AI, cs.CL, cs.LG
Submitted: 2026-03-09
Updated: 2026-10-08
Project page: https://attention-is-all-i-need.github.io/ACT
Importance score: 92/100
The gist: The gist The proposed Agentic Critical Training (ACT) is a reinforcement learning paradigm that trains large language models to autonomously develop reasoning about action quality by rewarding
Key concepts
- Agentic Critical Training (ACT)
- ACT is a reinforcement learning paradigm where agents are trained not just to mimic experts but to actively identify superior actions among choices. The model learns by being rewarded only when it correctly selects the better action, compelling it to develop its own reasoning about action quality.
- Imitation Learning (IL)
- Imitation learning involves training models by having them copy expert actions directly. This method teaches the agent what to do but fails because the model does not understand why an action is good or bad, resulting in a lack of awareness regarding action quality.
- Genuine Self-Reflection
- This refers to the model developing its own internal critical thinking process about its actions, rather than simply imitating pre-constructed reflection text. ACT achieves this by rewarding correct judgments, forcing the model to autonomously reason about why one action is superior to another.
- Contrastive Pairs
- These are pairs created during training where an expert action is contrasted with a model-generated alternative. By presenting these choices, the model learns to discriminate between actions, which is central to ACT's mechanism for developing critical reasoning.
Terminology
Summary
The gist The proposed Agentic Critical Training (ACT) is a reinforcement learning paradigm that trains large language models to autonomously develop reasoning about action quality by rewarding correct action selection, leading to genuine self-reflection rather than imitation of pre-constructed reflection text.
Agentic Critical Training
This approach moves beyond imitation learning by training agents to identify the better action among alternatives via reinforcement learning (RL) <ref:2603.08706#pg3>. The core idea is to transform the objective from “imitate the expert action” to “identify the better action,” requiring the model to develop discriminative understanding of action quality <ref:2603.08706#pg5>. ACT pairs each expert action with a model-generated alternative action to form a preference pair at each time step of sequential decision-making <ref:2603.08706#pg6>. The only supervision is whether the model correctly identifies the superior expert action, which drives the model to autonomously develop chain-of-thought (CoT; Wei et al., 2022) reasoning that leads to correct choices <ref:2603.08706#pg6>.
Training Pipeline and Reward Design
The training pipeline consists of two sequential RL stages: Agentic Critical Training followed by RL Action Training, both optimized using Group Relative Policy Optimization (GRPO; Shao et al., 2024) <ref:2603.08706#pg5>. The ACT prompt asks the model to decide which action is better among two candidates, and the model must output the chosen action inside tags <ref:2603.08706#pg19>. The composite reward function is defined as R(s, y) = Racc(a, a+) + Radm(a, Aadmissible) + Rfmt(y) <ref:2603.08706#pg16>. This function includes an accuracy reward for matching the expert action and an admissible action reward for valid but suboptimal actions <ref:2603.08706#pg16>.
Performance Across Benchmarks
ACT consistently improves agent performance across three challenging agent benchmarks (ALFWorld, WebShop, ScienceWorld) when combined with different post-training methods <ref:2603.08706#pg3>. Compared to imitation learning (IL), ACT achieves an average performance gain of 5.07 points, while outperforming reinforcement learning (RL) by 4.62 points <ref:2603.08706#pg9>. Furthermore, ACT demonstrates clear advantages over the Early Experience baseline, yielding an additional 2.42 points improvement on average <ref:2603.08706#pg6>.
Generalization and Reasoning Capabilities
ACT exhibits strong out-of-distribution generalization on agentic benchmarks and improves performance on general reasoning benchmarks (MATH-500 and GPQA-Diamond) without requiring any reasoning-specific training data <ref:2603.08706#pg4>. While IL and Early Experience fail to improve general reasoning, ACT achieves the highest scores on both MATH-500 and GPQA-Diamond despite being trained exclusively on agentic data <ref:2603.08706#pg12>. This is demonstrated by ACT exhibiting self-verification behavior on GPQA-Diamond, where it substitutes back into equations to check consistency <ref:2603.08706#pg12>.
Failure Recovery and Data Transferability
ACT enables failure recovery, as shown in a case study on ALFWorld where the IL model enters an infinite loop repeating a failed action for over 30 steps, whereas the ACT-trained model diagnoses the root cause and issues the correct navigation command <ref:2603.08706#pg10>. Moreover, ACT requires collecting alternative actions from a policy to construct contrastive pairs, but this data can be reused across model sizes for transferability <ref:2603.08706#pg12>. The results validate that ACT’s benefits generalize across model sizes and that the data collection cost can be amortized by reusing data across models of different sizes <ref:2603.08706#pg10>.
Conclusion
ACT trains LLM agents to reason about action quality by contrasting expert and self-generated actions via RL, producing autonomous critical reasoning through RL rather than imitating pre-constructed reflections <ref:2603.08706#pg12>. ACT not only enables strong out-of-distribution generalization on agentic benchmarks but also achieves notable improvements on general reasoning benchmarks without any reasoning-specific training data <ref:2603.08706#pg4>. ACT not only avoids the catastrophic forgetting observed in IL but improves upon the original model, suggesting that agentic RL environments may serve as a viable pathway for enhancing general reasoning capabilities <ref:2603.08706#pg12>. The results indicate that directly training models to evaluate action quality is more effective than supervising them to imitate reflection behaviors <ref:2603.08706#pg6>. ACT also improves OOD generalization, with the gain on top of RL being larger on OOD tasks (3.73pp) than on in-distribution tasks (2.15pp) <ref:2603.08706#pg9>. The paper concludes that ACT is a promising path toward developing more reflective and capable LLM agents <ref:2603.08706#pg2>.
--- Page 1 ---
Agentic Critical Training
Weize Liu, Minghui Liu, Sy-Tuyen Ho, Souradip Chakraborty, Xiyao Wang‡, Furong Huang‡ University of Maryland, College Park ‡Equal advising <ref:2603.08706#pg2>. Training large language models (LLMs) as autonomous agents often begins with imitation learning, but it only teaches agents what to do without understanding why: agents never contrast successful actions against suboptimal alternatives and thus lack awareness of action quality <ref:2603.08706#pg2>. Recent approaches attempt to address this by introducing self-reflection supervision derived from contrasts between expert and alternative actions <ref:2603.08706#pg2>. However, the training paradigm fundamentally remains imitation learning: the model imitates pre-constructed reflection text rather than learning to reason autonomously <ref:2603.08706#pg2>. We propose Agentic Critical Training (ACT), a reinforcement learning paradigm that trains agents to identify the better action among alternatives <ref:2603.08706#pg2>. By rewarding whether the model’s judgment is correct, ACT drives the model to autonomously develop reasoning about action quality, producing genuine self-reflection rather than imitating it <ref:2603.08706#pg2>. Across three challenging agent benchmarks, ACT consistently improves agent performance when combined with different post-training methods <ref:2603.08706#pg2>. It achieves an average improvement of 5.07 points over imitation learning and 4.62 points over reinforcement learning <ref:2603.08706#pg2>. Compared to approaches that inject reflection capability through knowledge distillation, ACT also demonstrates clear advantages, yielding an average improvement of 2.42 points <ref:2603.08706#pg2>. Moreover, ACT enables strong out-ofdistribution generalization on agentic benchmarks and improves performance on general reasoning benchmarks without any reasoning-specific training data <ref:2603.08706#pg2>. These results suggest that ACT is a promising path toward developing more reflective and capable LLM agents <ref:2603.08706#pg2>.
--- Page 2 ---
Agentic Critical Training
(a) Early Experience: Imitated Self-Reflection (b) ACT: Genuine Self-Reflection LLM Generates Reflection Action a∗ is better is better because it moves the agent closer to the target location, while Action a′ fails since the agent is not at the cabinet yet… Expert Action (a∗) Alternative Action (a′) Environment Next State (s∗) Next State (s′) Model Prompt Box Which action is better for the task? Action 1: put cloth in cabinet Action 2: go to cabinet Model Wait… I just cleaned the cloth at the sink. I need to go to the cabinet first before I can put it there! SFT Reflection is imitated from a fixed target string RL Reflection emerges autonomously through RL Figure 1: Comparison of imitated vs. genuine self-reflection. (a) Early Experience executes both actions in the environment, generates a reflection from the resulting states, and trains the model to imitate this fixed text via supervised fine-tuning (SFT) <ref:2603.08706#pg3>. (b) ACT presents two candidate actions and trains the model via RL to select the better one. Since only the selection outcome is rewarded, the model must autonomously develop reasoning about action quality to maximize reward <ref:2603.08706#pg3>.
Improvements for AI systems
-
Agentic Critical Training (ACT) drives autonomous reasoning by rewarding action quality judgment,
producing genuine self-reflection rather than imitating it.
This allows agents toautonomously develop chain-of-thought (CoT; Wei et al., 2022) reasoning that leads to correct choices
rather than imitating pre-constructed reflections. -
The improved system exhibits superior failure recovery by diagnosing root causes, as demonstrated when the IL model
enters an infinite loop repeating a failed action for over 30 steps until termination,
whereas ACT's model issues thecorrect navigation command
after self-critique. -
The agent gains strong out-of-distribution generalization, as ACT improves performance on general reasoning benchmarks without requiring reasoning-specific training data, suggesting that
learning to evaluate and compare actions can serve as a general mechanism for enhancing reasoning and decision-making abilities in LLM agents.
-
The system overcomes rigid execution errors by internalizing state awareness; instead of following
a rigid script (search → click item → select attributes → buy),
the ACT agent evaluates candidate actions against the current state, enabling it toback out and search again rather than blindly proceeding.
-
The model preserves and strengthens original reasoning capabilities on general benchmarks like MATH-500, as RL with ACT
avoids reasoning collapse because it optimizes for outcome correctness via RL rather than imitating behavioral patterns,
successfully avoiding the degradation seen in IL.
Sources
- GPT-4 Technical Report
- FireAct: Toward Language Agent Fine-tuning
- RM-R1: Reward Modeling as Reasoning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Mathematical Problem Solving With the MATH Dataset
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- Qwen3 Technical Report
- Agent Learning via Early Experience
- R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection