Knowledge-Based Zero-Replay Debugging of Multi-Agent LLM Traces
cs.SE, cs.AI
Submitted: 2026-06-11
Updated: 2026-09-26
Terminology
Sources
- Interactive Debugging and Steering of Multi-Agent AI Systems
- WhatIf: Interactive Exploration of LLM-Powered Social Simulations for Policy Reasoning
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
- TRAIL: Trace Reasoning and Agentic Issue Localization
- Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems
- MASPrism: Lightweight Failure Attribution for Multi-Agent Systems Using Prefill-Stage Signals
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- ReAct: Synergizing Reasoning and Acting in Language Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Generative Agents: Interactive Simulacra of Human Behavior
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies
- Training Verifiers to Solve Math Word Problems
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties