Facts Without Rules: Boundary Metadata Collapse in Multi-Agent LLM Handoffs
cs.AI
Submitted: 2026-08-29
Updated: 2026-08-29
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multi-agent LLM systems often coordinate by compressing an upstream interaction into a handoff artifact that downstream agents treat as shared state.
Abstract
Multi-agent LLM systems often coordinate by compressing an upstream interaction into a handoff artifact that downstream agents treat as shared state. We show that this handoff step is a structural source of privacy leakage: summaries preferentially preserve operational facts while weakening the boundary metadata that governs how those facts may be used---a failure mode we call summary collapse. On a controlled multi-agent coordination testbed we measure marker survival with a human-validated judge (κ= 0.74), where σ b = 1 means every boundary marker survives verbatim and σ b = 0 means all are lost. Boundary-marker and operational-fact survival are nearly uncorrelated at the handoff level on both GPT-5-mini and DeepSeek-R1-32B (Pearson r near zero): uncompressed free-text handoffs preserve boundaries at σ b about 0.80, whereas a 25-word budget drops σ b to about 0.57 while operational-fact survival stays near ceiling. Controlled downstream tests reveal that protection depends on boundary explicitness: vague languages leak in 73% of GPT and 50% of DeepSeek cases, while explicit constraints reduce leakage to under 15% across all three tested models. A no-handoff single-agent control further shows the failure is not reducible to multi-agent topology as direct full-marker access still leaks more often than the operationalized handoff. Prompt-only mitigation and exact-string redaction only partially address the problem, while a gold-derived audience allowlist nearly eliminates leakage across models, showing that correctly identifying audience boundaries is the key factor.
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection