StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems
Yanwen Peng, Delvin Ce Zhang, Xi Wang, Nikolaos Aletras
University of Sheffield
cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 18 pages, 3 figures, 4 tables, accepted by COLM2026
Code: https://github.com/YanwenPneg/StateBridge
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: StateBridge is a training-free latent communication approach for large language model (LLM) multi-agent systems.
Terminology
Summary
StateBridge is a training-free latent communication approach for large language model (LLM) multi-agent systems. It addresses the information loss that occurs when agents communicate through discrete text tokens, which discards continuous information from hidden states that token identities alone cannot capture. StateBridge aligns the sender’s final-layer hidden states to the receiver’s input embedding space via a closed-form orthogonal transformation, requiring no training and no architectural modification.
The method consists of three main components: message extraction, alignment interface, and prefix injection. For message extraction, the sender retains only the final K hidden states (e.g., K=64) of the generated message tokens, discarding intermediate reasoning sections. The alignment interface uses three steps: Procrustes alignment, norm calibration, and vocabulary anchoring. Procrustes alignment centers and whitens both the message hidden states and the reference embeddings (the embeddings of the corresponding decoded tokens), then solves an orthogonal Procrustes problem via singular value decomposition to find a rotation that preserves pairwise geometric relationships. Norm calibration rescales each aligned vector to the average l2 norm of the vocabulary embeddings, since final-layer hidden states have much larger norms than token embeddings (approximately 140× on Qwen3-4B). Vocabulary anchoring moves each calibrated vector slightly (with coefficient α=0.3) toward its nearest vocabulary embedding under cosine similarity, keeping the representation continuous while ensuring compatibility with the pretrained input distribution. The aligned prefix is then prepended to the receiver’s prompt embeddings, with position indices assigned sequentially, so the receiver processes it identically to ordinary token embeddings.
StateBridge is evaluated on eight benchmarks across three task categories: mathematical reasoning (GSM8K, AIME24, AIME25), question answering (GPQA-Diamond, ARC-Challenge, MedQA), and code generation (MBPP+, HumanEval+). Four models from two families are tested: Qwen3 (4B, 8B, 32B) and OLMo3-7B-Think. All agents share the same model weights within each run, using a four-agent sequential pipeline (Planner, Critic, Refiner, Judger). StateBridge is compared against three baselines: Single (standard single-agent generation), TextMAS (text-based sequential communication), and LatentMAS (KV-cache transfer).
Results show StateBridge achieves the best or tied-best score on 22 out of 26 model–task pairs, with average improvements of 2.4 to 2.9 points over the strongest baseline across all four model settings. On OLMo3-7B-Think, LatentMAS averages only 55.1% compared to TextMAS’s 73.9%, while StateBridge reaches 76.7%, demonstrating greater applicability across model families because it operates only at the input embedding layer rather than injecting states across all transformer layers. StateBridge helps most on challenging benchmarks including MedQA, GPQA, AIME24/25, and code generation tasks. On GSM8K for Qwen3 models, StateBridge trails the best baseline, likely due to exact-match evaluation sensitivity to output formatting.
Ablation studies isolate each component’s contribution. Replacing Procrustes with ridge regression (unconstrained linear map) drops average performance from 82.4% to 74.9%, with especially large drops on code generation tasks, because ridge regression distorts pairwise geometry and confines the prefix to the span of the discrete token embeddings. Removing norm calibration lowers the average to 79.5%, and removing vocabulary anchoring lowers it to 80.2%. Replacing the aligned prefix with random noise drops the average to 48.8%, ruling out gains from simply prepending extra continuous vectors.
Hyperparameter sensitivity analysis shows moderate values work best: increasing prefix length K from 16 to 64 usually improves performance, but K=128 hurts on most tasks because a single global rotation must compromise across a more heterogeneous point set. The anchoring coefficient α=0.3 strikes a consistent balance between informativeness and compatibility across all benchmarks.
Alignment visualization using PCA shows that before alignment, message hidden states occupy a different region from reference embeddings, but after alignment, the prefix moves much closer to the input embedding space, stable across model scales. A case study on a MedQA example demonstrates that the aligned prefix carries semantic information beyond the visible suffix tokens: at K=16, the critic recovers the diagnosis, key clinical features, and exclusion of alternative diagnoses, including the term “koilonychia” which does not appear in the visible suffix tokens. At K=64, the recovered plan becomes more detailed with differential diagnoses identified. At K=128, the critic copies only a truncated, incomplete plan, consistent with reduced alignment quality.
The paper provides theoretical analysis showing that text communication carries at most K log2 V bits (roughly 1101 bits for K=64 and V≈1.5×10 5), while continuous transmission is not subject to this combinatorial bound. It proves that ridge alignment confines the prefix to the span of the token embeddings (Proposition A.1), while orthogonal Procrustes preserves the Gram matrix and pairwise distances/angles of the whitened sender states (Proposition A.3).
Computational cost analysis shows the alignment is dominated by whitening eigendecomposition (O(d3)) and vocabulary anchoring search (O(KVd)), both cheaper than a single autoregressive pass for typical configurations. Space complexity is reduced to O(Kd) by using a forward hook on the final transformer layer, compared to O(TLd) for KV-cache transfer methods.
The paper concludes that the challenge in latent communication is not only sending richer states but making those states readable to the next agent, and that communication interface design matters in homogeneous multi-agent systems. Future work could extend the method to heterogeneous models, choose prefix length adaptively, and combine training-free alignment with learned communication modules.
Improvements for AI systems
Improvements to AI Systems:
- Cross-Model Latent Communication Without Fine-Tuning
-
Enable any LLM (e.g., Qwen, OLMo, Llama) to exchange continuous semantic states via a plug-in orthogonal alignment layer, eliminating the need for shared KV-cache or joint training.
-
The improved system can form heterogeneous multi-agent teams (e.g., a Qwen planner + OLMo critic) that collaborate on tasks while preserving each model’s pretrained capabilities.
- Lossless Information Transfer for Complex Reasoning
-
Replace discrete text summaries with aligned hidden-state prefixes (e.g., 64 vectors) that carry geometric relationships (distances/angles) beyond token identities.
-
The improved system can solve multi-step math (AIME), medical diagnosis (MedQA), and code generation with fewer communication rounds, recovering nuanced details (e.g., rare clinical terms like “koilonychia”) that text drops.
- Training-Free Adaptation to New Model Families
-
Apply the closed-form Procrustes + norm calibration + vocabulary anchoring pipeline to any new LLM in minutes, without gradient updates.
-
The improved system can be deployed on novel or proprietary models (e.g., future releases) with zero training data, achieving consistent gains over text-based baselines (e.g., +2.8% on OLMo3-7B).
- Adaptive Prefix Length Control
-
Dynamically select K (e.g., 16–64) based on task complexity and model scale, avoiding performance drops from over-long prefixes (K=128) that distort alignment.
-
The improved system can trade off communication bandwidth vs. accuracy: short prefixes for latency-sensitive tasks (e.g., chat), longer ones for high-stakes reasoning.
- Robustness to Output Formatting and Exact-Match Metrics
-
Mitigate text-based evaluation pitfalls (e.g., GSM8K format sensitivity) by injecting aligned latent context that guides the receiver’s generation without altering its decoding policy.
-
The improved system can maintain high performance even when the final answer format varies, reducing the need for post-hoc parsing or prompt engineering.
- Memory-Efficient Multi-Agent Pipelines
-
Use only O(Kd) memory for communication (vs. O(TLd) for KV-cache), enabling deployment on edge devices or long-horizon tasks with many agents.
-
The improved system can run 4+ agent pipelines (Planner→Critic→Refiner→Judger) on a single GPU with lower VRAM footprint, scaling to larger models or longer contexts.
- Semantic Consistency Across Agent Roles
-
Preserve pairwise geometric structure of hidden states via orthogonal rotation, ensuring that the receiver’s internal representations remain coherent when processing the prefix.
-
The improved system can maintain role-specific behaviors (e.g., critic’s strict evaluation, refiner’s targeted edits) without catastrophic forgetting or mode collapse.
- Interpretable Latent Communication
-
Visualize aligned prefixes via PCA to verify that they land near the input embedding space, enabling debugging of multi-agent failures.
-
The improved system can provide human-readable explanations of what latent information was transferred (e.g., “critic received evidence for differential diagnosis X”), improving trust and auditability.
Abstract
Large language model based multi-agent systems usually communicate in text, i.e., using discrete tokens. However, text introduces a discrete bottleneck. Converting the sender's continuous hidden states into discrete tokens discards information that token identities alone cannot capture. Recent work proposes latent communication as an alternative, where agents transmit hidden representations directly without converting them to text. However, existing latent methods either inject working memory layer by layer across the transformers, or require trained projectors that limit portability. We propose StateBridge, a training-free latent communication approach that aligns the sender's final-layer hidden states to the receiver's input space via a closed-form orthogonal transformation. Lightweight norm calibration and vocabulary anchoring ensure compatibility with the pretrained input distribution. The aligned states are prepended to the input of the receiver agent as a continuous prefix. We evaluate StateBridge on math reasoning, code generation, and question answering with four models from two families. StateBridge achieves the best or tied-best score on 22 out of 26 model-task pairs, consistently outperforming the strongest baseline.
Sources
- Optimal Transport Depth Up-Scaling
- Evaluating Large Language Models Trained on Code
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks
- Predicting Multi-Agent Specialization via Task Parallelizability
- Olmo 3
- Talk Structurally, Act Hierarchically: A Collaborative Framework for LLM Multi-Agent Systems
- Talk to Right Specialists: Iterative Routing in Multi-agent Systems for Question Answering
- Beyond Self-Talk: A Communication-Centric Survey of LLM-Based Multi-Agent Systems
- Qwen3 Technical Report
- LLM-based Multi-Agent Systems: Techniques and Business Perspectives
- Mobile-Agent-v3: Fundamental Agents for GUI Automation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection