Locating Hidden Failures Makes Long-Horizon Agents More Reliable
cs.LG
Submitted: 2026-09-15
Updated: 2026-09-15
License: http://creativecommons.org/licenses/by/4.0/
The gist: As AI agents take on long, autonomous tasks, we increasingly oversee rather than perform the work, yet we still judge them almost entirely by whether they finally succeed.
Terminology
Abstract
As AI agents take on long, autonomous tasks, we increasingly oversee rather than perform the work, yet we still judge them almost entirely by whether they finally succeed. An outcome cannot reveal where a run went wrong, whether the agent recovered, or the irreversible harm it caused along the way, and where long-horizon agents fail remains unmapped. We study 2518 agent trajectories across software engineering, computer use, and science, close to real deployment, and classify 6967 mistakes into 78 failure types. Failure follows a recurring signature: after its first mistake an agent often fails to recover and rarely catches the error itself, so the run continues unchecked while still looking correct; whether an agent recovers depends on the task and the environment's feedback, not on the agent framework running it. Long-horizon agents can do real harm on the way to a passing result: even runs scored as solved delete data, corrupt systems, or fabricate success rather than earning it. We release these human-verified annotations as Traverse, a benchmark on which six frontier judges struggle to locate failure regardless of scale: even the strongest correctly identifies the first mistake in fewer than a third of runs. Yet Scout, a 4 B verifier we trained, locates failure far better than these judges and transfers to domains it never saw. Used at test time to select among an agent's candidate runs, it raises task success above the agent's own single-attempt performance, without retraining the agent. By making failure cheap to locate and correct, this work is a foundation for more trustworthy long-horizon agents that learn from their own mistakes, and a practical path to overseeing increasingly autonomous AI.
Sources
- Measuring Progress on Scalable Oversight for Large Language Models
- CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- TRAIL: Trace Reasoning and Agentic Issue Localization
- GLM-5: from Vibe Coding to Agentic Engineering
- Alignment faking in large language models
- ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
- CodeTracer: Towards Traceable Agent States
- Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers
- AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
- ScalePRM: Training Process Reward Models by Scaling Verification Compute Without Ground Truth
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents
- Large Language Models can Strategically Deceive their Users when Put Under Pressure
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
- Agents' Last Exam
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks