Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework
Jing Chen, Yang Sun, Li Zhang, Lin Xu, Jie Shi
cs.AI
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: 17 pages, 4 figures, 7 tables
Code: https://github.com/chenjing-2024/agent-trajectory-attribution
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
- Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
- Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation
- CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution
- AgentCanary: A Security Evaluation Framework for Autonomous AI Agents in Real Executable Environments
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
- AgentBench: Evaluating LLMs as Agents
- ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents
- The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution
- Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures
- HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
- MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection