Faithful, Not Corrective: Model Capability Governs Message-Format Effects in Multi-Hop Agent Relays
cs.AI, cs.LG
Submitted: 2026-06-12
Updated: 2026-09-17
License: http://creativecommons.org/licenses/by/4.0/
The gist: When LLM agents hand information to one another, does the message format matter? Two literatures disagree: format-optimization work reports that structured messages cut cost without hurting accuracy,
Terminology
Abstract
When LLM agents hand information to one another, does the message format matter? Two literatures disagree: format-optimization work reports that structured messages cut cost without hurting accuracy, while format-restriction studies find that imposing structure degrades generation. Neither line has measured what happens when messages traverse multiple hops, where copy fidelity, rather than one-shot generation quality, dominates. We introduce a controlled relay testbed in which briefs of twelve programmatic atomic facts are re-encoded hop by hop in five formats (free natural language, precision-instructed NL, JSON, triples, key-value) over six hops, scored against programmatic ground truth by a fixed strong grader, across two relay-capability tiers, a cognitive-load condition, and a paired-fork error injection. We find that (i) a strong relay is nearly lossless for every format (hop-6 QA recall at least 0.973), with residual loss concentrated at the first encoding step; (ii) per-hop cognitive load raises generation cost by 24-53% while fidelity changes stay within plus or minus 1.8 points; (iii) under a weak 1.5B relay, the across-format dispersion of hop-6 recall grows by a factor of 8.7 (CI 5.3-15.5), driven by an encode-drift trade-off that flips the format ranking in transit; and (iv) once an injected error is present, every format propagates it faithfully (surface persistence 83-100%) and no format cascades collateral damage onto neighboring facts. Structure buys a faithful, error-localizing channel, not an error-correcting code.
Sources
- On the Reliability Limits of LLM-Based Multi-Agent Planning
- Why Do Multi-Agent LLM Systems Fail?
- A survey of agent interoperability protocols: Model Context Protocol (MCP), Agent Communication Protocol (ACP), Agent-to-Agent Protocol (A2A), and Agent Network Protocol (ANP)
- Information Fidelity in Tool-Using LLM Agents: A Martingale Analysis of the Model Context Protocol
- Lost Before Translation: Social Information Transmission and Survival in AI-AI Communication
- Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents
- Towards a Science of Scaling Agent Systems
- A Scalable Communication Protocol for Networks of Large Language Models
- When LLMs Play the Telephone Game: Cultural Attractors as Conceptual Tools to Evaluate LLMs in Multi-turn Settings
- Qwen2.5 Technical Report
- Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
- Talk Structurally, Act Hierarchically: A Collaborative Framework for LLM Multi-Agent Systems
- Toward a Safe Internet of Agents
- From Spark to Fire: Modeling and Mitigating Error Cascades in LLM-Based Multi-Agent Collaboration
- The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
- Can Small Agents Collaborate to Beat a Single Large Language Model?
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection