Chronicle: Cut-Point Replay for Regression Testing of LLM Agents
cs.CL, cs.AI
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/theagentplane/chronicle
Project page: https://langchain-ai.github.io/langgraph
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing
Terminology
Abstract
Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step trajectory that a re-run rarely repeats. Record-and-replay makes a run reproducible, but existing agent tooling records runs only to trace or score them, not to test a code change against them. We present Chronicle, which records an agent run at its non-deterministic boundaries as immutable envelopes and replays it from the record. Its central operation, cut-point replay, serves a chosen subset of boundaries from the record and executes the complementary subset live with new code, turning a recorded incident into a regression test that runs in continuous integration. On a benchmark of 6 recorded failures with simulated model boundaries, recording adds 23 μs per crossing (0.008% of an assumed 300 ms model call), full replay issues zero model calls and is bit-stable across 20 repetitions, and cut-point tests fail on faulty code and pass on guarded and benign changes for all 6 incidents. In a mutation study of the guarded tools, cut-point tests catch every mutant that lets the recorded unsafe action through, while a baseline that stubs every boundary, using the same assertion, catches none. Chronicle and the benchmark are publicly available at https://github.com/theagentplane/chronicle.
Sources
- Why Do Multi-Agent LLM Systems Fail?
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- TRAIL: Trace Reasoning and Agentic Issue Localization
- Get Experience from Practice: LLM Agents with Record & Replay
- Detecting and Evaluating Order-Dependent Flaky Tests in JavaScript
- Are Coding Agents Generating Over-Mocked Tests? An Empirical Study
- Knowledge-Based Zero-Replay Debugging of Multi-Agent LLM Traces
- Automated structural testing of LLM-based agents: methods, framework, and case studies
- REFLECT: Intervention-Supported Error Attribution for Silent Failures in LLM Agent Traces
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
- Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias
- Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering