Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions
cs.AI
Submitted: 2026-09-02
Updated: 2026-09-02
Code: https://github.com/langchain-ai/langchain
License: http://creativecommons.org/licenses/by/4.0/
The gist: With the proliferation of LLM agents, the ability to understand and diagnose failures in agents is essential to achieving superior effectiveness and trustworthiness.
Terminology
Abstract
With the proliferation of LLM agents, the ability to understand and diagnose failures in agents is essential to achieving superior effectiveness and trustworthiness. As agent failures often manifest via long and complex trajectories, manually finding the needles in the haystack is untenable. However, traditional diagnosis techniques for software bugs can hardly address LLM agent failures, while completely relying on LLMs as the judge yields unreliable diagnosis results. To overcome these challenges, this paper presents AGENTSCOPE, a new neuro-symbolic approach for agent failure mode diagnosis. The key principle of AGENTSCOPE is to abstract agent behavior, based on its trajectories, into structured representations. Furthermore, AGENTSCOPE introduces the concept of neural invariants to specify agent behavior properties. AGENTSCOPE leverages LLM-guided reasoning atop the structured representation against neural invariants to pinpoint both the failure step and its type in the trajectory. We show the effectiveness of AGENTSCOPE on publicly available agent failure datasets (Who&When) and a more comprehensive dataset created by us (AgentErrata), where AGENTSCOPE significantly outperforms the current state of the art in fault localization and attribution accuracy. Our work shows that integrating structured abstractions with LLM-guided reasoning enables effective, reliable, and interpretable diagnosis for agent failures.
Sources
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Why Do Multi-Agent LLM Systems Fail?
- Humans or LLMs as the Judge? A Study on Judgement Biases
- Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
- AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn
- A Survey on LLM-as-a-Judge
- OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
- GPT-4o System Card
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
- DeepSeek-V3 Technical Report
- OpenAI GPT-5 System Card
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Qwen3 Technical Report
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- RAFFLES: Reasoning-based Attribution of Faults for LLM Systems
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection