CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents
cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted to EMNLP 2026 (Main)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) agents are increasingly deployed in long-horizon, interactive, and stateful environments.
Terminology
Abstract
Large language model (LLM) agents are increasingly deployed in long-horizon, interactive, and stateful environments. In these settings, a single wrong action, such as refunding the wrong purchase, can cause irreversible task failure and must be intercepted before execution. Such failures may not appear in every single run, but can emerge across repeated trials, making reliability across steps and trials critical. However, ensuring agentic reliability is challenging: even frontier LLMs struggle to explain why an action may be wrong, especially in long, intertwined trajectories governed by domain-specific policies. Much recent work relies on prompt-based critique agents, while optimization-based methods lack a systematic way to produce rich verification rationales for training. We address this gap with CAST, a critique-aware training framework that converts sparse task outcomes into action-level supervision for critique learning and policy optimization. CAST analyzes agent trajectories to synthesize structured rationales explaining action validity under partial observability. The resulting critique model is used to construct critique-aware training data for optimizing the policy model. Fine-tuning Qwen3-family models on dynamic tool-calling benchmarks, CAST improves reliability across domains, outperforming GPT-OSS-120B by over 10% pass 4 on Retail tasks and yielding an additional 9% improvement on Telehealth in an out-of-domain setting. These results demonstrate that critique-aware training improves the robustness of LLM agents in realistic dynamic environments.
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
- Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
- Asymmetric Actor-Critic for Multi-turn LLM Agents
- LLMs Get Lost In Multi-Turn Conversation
- Dancing with Critiques: Enhancing LLM Reasoning with Stepwise Natural Language Self-Critique
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
- Self-Refine: Iterative Refinement with Self-Feedback
- How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench
- GPT-4 Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
- Self-Reflection in LLM Agents: Effects on Problem-Solving Performance
- FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
- LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents
- PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
- Qwen3 Technical Report
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering