TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories
Yunjia Qi, Zehua Yin, Xintong Shi, Hao Peng, Songyuanyi Lu, Yixian Liu, Richeng Xuan, Yuhong Liu, Zhichao Hu, Xiaozhi Wang, Lei Hou, Bin Xu, Juanzi Li
Department of Computer Science and Technology, Beijing National Research Center for Information Science and Technology, Tsinghua University · Shenzhen International Graduate School, Tsinghua University · Tencent Hunyuan
cs.AI
Submitted: 2026-08-06
Code: https://github.com/THU-KEG/TrajDebug
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 68/100
The gist: The paper addresses the problem of critical error detection in long-horizon LLM-based agent trajectories.
Terminology
Summary
The paper addresses the problem of critical error detection in long-horizon LLM-based agent trajectories. As the abstract states: "LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that is responsible for the final failure. The paper identifies two main challenges: (1) long trajectories make individual error identification difficult because
the evidence for judging a step may be scattered across distant instructions, observations, and prior context, and (2)
failed trajectories often contain multiple local errors whose downstream effects differ: some are repaired, some remain harmless, and only some contribute to the final failure."
The authors propose TRAJDEBUG, an error-lifecycle tracing framework
that addresses long-trajectory error discovery with multi-granularity history compression and evidence-based error identification, and supports critical attribution by tracing each error's resolution status and terminal impact.
They also construct TRAJERRBENCH, a benchmark of 486 manually annotated failed trajectories from τ2-Bench and SWE-Bench Pro, covering realistic tool-use and coding scenarios.
TRAJDEBUG is a three-stage framework that decomposes the task into three stages: error trigger detection, error state classification, and critical attribution.
Stage 1 — Multi-Granularity Compression (§3.1): The framework constructs three views for each step: (1) high-detail view keeps the original instruction, action, observation, and locally relevant reasoning snippets needed for evidence verification
; (2) medium-detail view summarizes the step's main intent, action, and state update
; (3) low-detail view records only coarse progress, salient entities, and unresolved commitments.
Each stage retrieves the finest view required for its decision, using high-detail views for local verification and compressed views for distant context.
Stage 2 — Error Trigger Detection (§3.2): The first stage extracts per-step atomic error triggers.
A trigger e = (t, c, p, qw, qr) denotes a local evidence-grounded mismatch at step t, where qw is an erroneous commitment expressed in the current step, qr is the violated reference, c is the reference category, and p is the execution phase.
Crucially, to prevent subjective drift and hallucinated diagnoses, each trigger must satisfy a verbatim evidence condition: both qw and qr must be explicitly citable; otherwise, the trigger is discarded.
Triggers are organized along two axes. The reference category c includes: (1) Task Conflict (qr from the task instruction), (2) History Conflict (qr from prior trajectory context), (3) Intra-Step Conflict (qr from the current step itself), and (4) Environment Anomaly (when the agent action is reasonable but the environment response is abnormal
). The execution phase p records where the mismatch surfaces: planning, reasoning, action, observation, or verification.
Stage 3 — Error State Classification (§3.3): Triggers are grouped into error instances
where an instance as E = (E, O), where E contains triggers that violate the same reference object O.
The state classifier makes two evidence-backed judgments: whether the wrong commitment is resolved or remains active, and whether it leaves an observable terminal footprint.
Combining these yields four error states:
-
Clean Resolution (resolved, no footprint)
-
Costly Resolution (resolved, budget-debt footprint)
-
Manifest Active (active, irreversible or semantic footprint)
-
Latent Active (active, no footprint)
Terminal-relevant candidates are retained: F(τ) = E: state(E) ∈ CostlyResolution, ManifestActive.
Stage 4 — Candidate-Set-Guided Causal Attribution (§3.4): This attribution head selects the candidate instance whose first step best explains the terminal failure.
A simple earliest-candidate rule is insufficient because temporal order alone does not determine failure responsibility.
The LLM attribution head is given each candidate's first step, state label, and supporting verbatim evidence.
Existing benchmarks remain limited in scale, domain diversity, or trajectory complexity.
TRAJERRBENCH contains 486 manually annotated failed trajectories from two complementary domains: 400 τ2-Bench trajectories and 86 SWE-Bench Pro trajectories, averaging 29.3 and 119.7 steps, respectively.
Each trajectory is annotated by three annotators with majority-vote labels. Agreement statistics: The critical-error step labels reach almost-perfect agreement on τ2-Bench (Fleiss' κ=0.91) and substantial agreement on the longer, code-heavy SWE-Bench Pro trajectories (Fleiss' κ=0.67).
The authors note that critical errors in TRAJERRBENCH often occur in the mid-to-late
portion of trajectories, reflecting that agents first gather task-relevant information before making decisive decisions.
In τ2-Bench, critical errors are almost evenly split between task conflicts (52.4%) and history conflicts (46.3%), while SWE-Bench Pro is dominated by task conflicts (73.4%).
Reasoning is the dominant failure phase in both domains (60.7% and 57.1%).
Main results (Table 1): TRAJDEBUG achieves the best overall average in our benchmark suite,
obtaining the highest macro-average accuracy (34.11%),
outperforming both direct prompting with frontier models and multi-agent baselines. It improves over direct prompting with the same backbone from 25.69% to 34.11% (+8.42).
TRAJDEBUG ranks first on five of seven datasets, spanning embodied decision-making, open-domain information seeking, user-facing tool use, and long-horizon software engineering.
On the longest-horizon benchmarks, TRAJDEBUG attains the best accuracy with gains of +13.00% and +6.97% over the corresponding direct baselines
on ALFWorld and SWE-Bench Pro, respectively.
Ablation study (Table 2): Replacing multi-granularity compression with the same hard-truncation strategy used by direct prompting causes the largest drop, reducing AVG by 21.02 points.
Removing error trigger detection or error state classification lowers AVG by 4.99 and 7.31 points, respectively.
Length analysis (Figure 5): "All baselines degrade sharply as trajectories grow, dropping from 35–50% on short trajectories to below 15% on long ones. In contrast, TRAJDEBUG remains substantially more stable, retaining over 20% accuracy on the longest bucket."
Application studies (Table 3): In per-trajectory repair under oracle failure, TRAJDEBUG delivers the largest gains across all three settings, lifting the average success rate from 78.07 to 88.87,
achieving a 10.80% average improvement.
In the failure-memory transfer scenario (detector applied to 1/3 of historical failures, feedback aggregated and injected when evaluating on the held-out 2/3 split), TRAJDEBUG still yields the largest improvement, yielding a 5.70% average improvement,
whereas Vanilla and Self-Reflection show less stable transfer and can even reduce performance in some cases.
The authors conclude: "We presented TRAJDEBUG, a three-stage error-lifecycle tracing framework for critical error detection. It combines error trigger detection, error state classification, and candidate-set-guided attribution to identify errors in long trajectories and distinguish failure-responsible errors among multiple errors... Experiments show that TRAJDEBUG outperforms LLM prompting and diagnostic baselines, and application studies demonstrate its value for trajectory repair and transferable failure memory. These results position critical error detection as a promising interface for inference-time agent improvement." The code and data will be released under the MIT license.
The paper acknowledges two main limitations: (1) Although TRAJDEBUG constrains judgments with verbatim evidence and structured candidates, it still relies on LLMs for error interpretation and final attribution,
and (2) As a staged pipeline, TRAJDEBUG may inherit false negatives from earlier trigger detection or state classification stages, causing the final candidate set to miss the true critical error.
The error analysis (§E.5) shows the failure decomposition: Trigger Misses (42.1%) and Attribution Misses (40.1%) are the two dominant failure modes, while State Misses are less frequent (17.7%).
Improvements for AI systems
-
Add multi-granularity trajectory memory for long-horizon agents. Instead of storing one flat transcript or hard-truncating old context, the agent stores each step in three detail levels: high-detail (original instruction, action, observation, verbatim reasoning snippets), medium-detail (step intent, action, state update), and low-detail (coarse progress, salient entities, unresolved commitments). When diagnosing failures, it retrieves the finest view needed for local evidence and compressed views for distant context. This prevents accuracy collapse on long trajectories.
-
Add evidence-grounded error trigger extraction with structured failure records. The agent detects atomic errors as tuples
(t, c, p, qw, qr)wheretis the step,cis the reference category (task conflict, history conflict, intra-step conflict, environment anomaly),pis the execution phase (planning, reasoning, action, observation, verification),qwis the erroneous commitment in the current step, andqris the violated reference. Crucially, it only emits a trigger if bothqwandqrcan be cited verbatim; otherwise it discards the trigger. This suppresses hallucinated diagnoses and makes every failure explanation auditable. -
Add error-state lifecycle classification. The agent groups triggers that violate the same reference object into error instances, then classifies each instance into one of four states:
-
Clean Resolution: the wrong commitment was fixed with no terminal footprint.
-
Costly Resolution: fixed, but left budget/step debt that contributes to failure.
-
Manifest Active: still active and leaves an irreversible or semantic footprint.
-
Latent Active: still active but leaves no observable footprint.
Only Costly Resolution and Manifest Active are retained as terminal-relevant candidates. This lets the system ignore repaired errors and harmless mistakes, and focus only on errors that actually cause the final failure.
-
Add candidate-set-guided causal attribution. Instead of selecting the earliest error, the system uses an LLM attribution head that receives each candidate’s first step, state label, and supporting verbatim evidence, and chooses the instance whose origin best explains the terminal failure. This avoids the false assumption that temporal order alone determines responsibility.
-
Add transferable failure memory. Store validated critical errors—including their evidence, state label, and attribution—and inject them as episodic feedback when similar situations appear in future trajectories. This gives the agent a growing memory of past failure patterns and improves performance on new, held-out tasks.
-
Add targeted repair scheduling. Once a Costly Resolution or Manifest Active error is identified, the agent can roll back to the responsible step or inject a corrective plan before continuing. This enables online recovery during long-horizon tasks rather than only post-hoc analysis.
-
Identify the critical root-cause error in long failed trajectories even when the evidence is scattered across distant instructions, observations, and prior context.
-
Maintain diagnostic accuracy as trajectory length grows. Where direct prompting and multi-agent baselines drop below 15% accuracy on the longest trajectories, the improved system remains substantially more stable and retains meaningful accuracy.
-
Distinguish fatal errors from benign or repaired errors. When a trajectory contains multiple mistakes, it correctly ignores ones that were fixed or harmless and focuses on the few that lead to the terminal failure.
-
Provide human-auditable explanations for every diagnosis. Each critical-error judgment is supported by exact verbatim quotes of the erroneous statement and the violated reference, so users can verify the reasoning and trust the model’s self-debugging.
-
Recover during task execution. If it detects a manifested critical error while still running, it can perform a targeted repair at the responsible step, raising task success substantially—on average from around 78% to 89% in the paper’s repair settings.
-
Transfer lessons from past failures to future tasks. With failure memory collected from only a third of historical failures, the system improves success on held-out tasks by about 5.7% on average, while vanilla self-reflection methods can actually reduce performance.
-
Work across diverse long-horizon domains: tool-use benchmarks (e.g., τ2-Bench), software engineering (e.g., SWE-Bench Pro), embodied decision-making, and open-domain information seeking—because the error representation and evidence constraints are domain-agnostic.
-
Reduce hallucinated self-diagnoses. The verbatim-evidence requirement means the system refuses to generate speculative error explanations when no concrete contradiction exists, making it a more reliable inference-time interface for agent debugging and improvement.
Sources
- GPT-4 Technical Report
- Where Did It All Go Wrong? A Hierarchical Look into Multi-Agent Error Attribution
- AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- Demystifying the Lifecycle of Failures in Platform-Orchestrated Agentic Workflows
- TRAIL: Trace Reasoning and Agentic Issue Localization
- AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents
- AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems
- When does In-context Learning Fall Short and Why? A Study on Specification-Heavy Tasks
- Who is Introducing the Failure? Automatically Attributing Failures of Multi-Agent Systems via Spectrum Analysis
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Self-Reflection in LLM Agents: Effects on Problem-Solving Performance
- Gemini: A Family of Highly Capable Multimodal Models
- Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation
- Kimi K2.5: Visual Agentic Intelligence
- CodeTracer: Towards Traceable Agent States
- From Flat Logs to Causal Graphs: Hierarchical Failure Attribution for LLM-based Multi-Agent Systems
- AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection