Stop Means Stop: Measuring and Repairing the Enforcement Gap in Agent-Framework Control Primitives
Sajjad Khan
cs.SE, cs.CR, cs.DC
Submitted: 2026-07-15
Comments: 32 pages, 3 figures, 11 tables. Code: pip install soundgate (PyPI)
Code: https://github.com/sajjadanwar0/soundgate-paper
License: http://creativecommons.org/licenses/by/4.0/
The gist: Production LLM-agent frameworks ship control primitives -- human-in-the-loop approval gates, run cancellation, and execution timeouts -- whose names and documentation imply barrier semantics: while a
Terminology
Abstract
Production LLM-agent frameworks ship control primitives -- human-in-the-loop approval gates, run cancellation, and execution timeouts -- whose names and documentation imply barrier semantics: while a run is paused, cancelled, or timed out, no gated side effect executes. This contract holds on none of six widely used open-source frameworks. Model-free differential probes isolate a recurring sibling leak -- an approval gate suspends its own branch while a sibling's effect executes during the pause, defeating rejection -- in every framework shipping a pre-execution gate (five of six, four execution models, two language runtimes), and confirm replay double-execution, cancellation orphans, and timeout zombies. The hazard is reachable: frontier models emit the leak-triggering plan shape at rates up to 14%, and live models driving unmodified frameworks leak 215 of 1,200 runs (P(leak emitted)=1.00); on naturalistic tau-bench episodes models serialize writes -- the everyday gap is latent -- while injection induces it deterministically and a 13-incident public corpus corroborates the replay and cancellation failures. We repair the gaps with SOUNDGATE, an environment-external Rust gate through which every side effect must be admitted, enforcing hold-until-decided, reject-cancels, dedup-on-replay, and fence-on-cancel under a stated complete-mediation contract, discharged for network egress by two kernel-enforced routes. The admission core is mechanically verified (Verus; TLA+/TLC to 7.5e7 states; TLAPS; Loom on the deployed Rust) and bridged to code by differential conformance over 1.2e7 operations with zero divergences. Under that contract SOUNDGATE blocks every measured violation on all six frameworks while releasing legitimate effects: gated tau-bench episodes complete with zero refusals at 1 ms per write, and durable admission sustains 12k admissions per second.
Sources
- AI Control: Improving Safety Despite Intentional Subversion
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
- Progent: Securing AI Agents with Privilege Control
- Defeating Prompt Injections by Design
- Mind the Gap: Time-of-Check to Time-of-Use Vulnerabilities in LLM-Enabled Agents
- Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents
- Why Do Multi-Agent LLM Systems Fail?
- Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent Large Language Model Systems
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- MemGPT: Towards LLMs as Operating Systems
- AIOS: LLM Agent Operating System
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties