Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations
cs.AI
Submitted: 2026-05-13
Updated: 2026-09-09
License: http://creativecommons.org/licenses/by/4.0/
The gist: In a long conversation, an LLM can produce a plausible continuation that rests on premises the conversation has already abandoned.
Terminology
Abstract
In a long conversation, an LLM can produce a plausible continuation that rests on premises the conversation has already abandoned. No runtime check ties its output to what the conversation has established, a gap that context-manipulation attacks on deployed agents exploit. We close this gap with a runtime verifier: an LLM Interpreter classifies each utterance into one of eight epistemic operations, and a symbolic engine applies them to a dependency map that records what every claim rests on and whether it still stands. Whether a continuation is grounded reduces to a walk over the map, linear in its size, with no LLM call. Retraction propagates through the same map with a conflict-free guarantee, flagging exactly the conclusions that lose support. On ReviseQA for belief revision and MemoryAgentBench's fact-consolidation split, two third-party benchmarks where earlier premises are superseded, the verifier leads a budget-matched retrieval baseline across five QA models and lifts MemoryAgentBench single-hop accuracy from 0.46--0.95 to 0.93--0.98. With the verifier, even the 7B model overtakes unaided GPT-4o. These runs feed the engine the benchmarks' own structured updates. When a GPT-4o Interpreter extracts every update from raw text instead, accuracy is statistically unchanged. Per-query cost is flat in conversation length, prompts staying near 0.8k tokens where full context reaches 114k and retraction queries under a microsecond at 2000 turns.
Sources
- Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models
- LLM-State: Open World State Representation for Long-horizon Task Planning with Large Language Model
- Memory Injection Attacks on LLM Agents via Query-Only Interaction
- D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
- Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory
- WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment
- AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
- Beyond the Black Box: Demystifying Multi-Turn LLM Reasoning with VISTA
- A Survey on the Memory Mechanism of Large Language Model based Agents
- WildChat: 1M ChatGPT Interaction Logs in the Wild
- WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection