Localizing Emergent Failures in Agentic AI: Recovering Minimal Repair Families via Counterfactual Replay
cs.AI, cs.MA
Submitted: 2026-08-29
Updated: 2026-09-06
Comments: 6 pages, conference paper
License: http://creativecommons.org/licenses/by/4.0/
The gist: Failures in agentic AI systems can arise from interactions among messages exchanged by multiple large language model (LLM) agents.
Terminology
Abstract
Failures in agentic AI systems can arise from interactions among messages exchanged by multiple large language model (LLM) agents. Pointwise attribution cannot distinguish a jointly necessary repair from alternative singleton repairs. We formulate Minimal Repair Family Recovery (MRFR): recovering all inclusion-minimal event sets whose counterfactual replay restores task success within a declared size bound. We propose Graph-Constrained Joint Replay (GCJR), which slices failure-relevant events from an execution dependency graph, constructs graph-feasible singleton and pair candidates, and verifies them by replay with paired clean counterparts. For fixed replay outcomes, GCJR is exact within its declared graph domain. On 90 in-scope cases from a 120-DAG controlled benchmark, GCJR achieves 1.000 Family Exact Match while reducing mean replay calls from 56.3 to 25.3 (55.1%) relative to exhaustive search. On a 24-case, four-agent LLM pilot, it again achieves 1.000 Family Exact Match and reduces mean model calls from 21.0 to 10.0 (52.4%); single-event replay misses jointly necessary repairs.
Sources
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
- GraphTracer: Graph-Guided Failure Tracing in LLM Agents for Robust Multi-Turn Deep Search
- From Flat Logs to Causal Graphs: Hierarchical Failure Attribution for LLM-based Multi-Agent Systems
- HPFA: Hypergraph-Based Paired Failure Attribution for LLM Reasoning
- CausalFlow: Causal Attribution and Counterfactual Repair for LLM Agent Failures
- Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures
- Spatio-Temporal Trajectory Similarity Measures: A Comprehensive Survey and Quantitative Study
- Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection