When Less Latent Leads to Better Relay: Information-Preserving Compression for Latent Multi-Agent LLM Collaboration

arXiv:2604.13349 · cs.LG · Submitted 2026-08-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "When Less Latent Leads to Better Relay: Information-Preserving Compression for Latent Multi-Agent LLM Collaboration".

Jane: The paper was written by Yiping Li, Zhiyu An and Wan Du from University of California, Merced and Department of Computer Science and Engineering at University of California, Merced (as part of the affiliation).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: We’ve seen the general idea of compression, but now let's look at what the results actually tell us about "When Less Latent Leads to Better Relay."

Jane: The authors show that compressed methods—L-OBF and H-OBF—are achieving accuracy levels that are remarkably close to or even surpassing what we consider the full KV state transfer baseline.

Lu: It’s a powerful demonstration that just having a lot of information doesn' not necessarily lead to better results; preserving the *most useful* information matters more.

Meng: The data shows L-OBF and H-OBF sit around one point seven to one point nine percentage points above Full on average across nine different benchmarks, which is a significant operational achievement.

Lalam: This suggests that we can build systems that are truly leveraging context without the massive state transfer, allowing us to manage our expectations of what's possible in AI design.

Tom: It’s interesting because this performance isn't just due to a lucky streak; it’s consistent with the authors demonstrating this level of performance across mathematical reasoning, expert QA, and coding tasks.

Jane: I think it’s reassuring to see that these improvements hold up across such diverse benchmarks; it isn't just working on one specific type of problem, it is consistently performing well across all nine categories.

Meng: This consistency in performance gives us confidence that we can scale this approach to handle large, multi-step tasks without worrying about a single point of failure due to excessive communication costs.

Lu: This opens up pathways for truly massive, complex multi-agent problems that were previously limited by the sheer volume of communication we had to send between agents.

Lalam: It shows our AI systems can be designed to be both incredibly powerful and efficient, making a huge leap forward in how we deploy these agents in the real world by minimizing overhead.

Paper discussion segment 2: Tom: Moving beyond the results, we need to talk about *why* this works—the core mechanism of "When Less Latent Leads to Better Relay."

Jane: The authors aren't just performing a simple hard eviction; instead of just discarding the lost information during compression, they are finding its unique components that don't overlap with what we kept, which is the residual content.

Lu: And then injecting those back into the retained values is an incredibly sophisticated way to manage information loss without causing double-counting errors in our model weights; it’s a very elegant solution.

Meng: The engineering challenge here is ensuring that this injection doesn't add too much corrective signal, and OBF handles that by calculating a dynamic scaling factor based on how much attention was originally paid to the deleted tokens, which is crucial for stability.

Lalam: This allows us to solve a fundamental problem in AI communication; we can now recover lost context precisely without disrupting the flow of information between agents, making our systems far more reliable and dependable over time.

Tom: That’s right; the paper shows that by applying this sophisticated backfill, the error bound is significantly reduced compared to just letting the data go, providing a measurable improvement in overall fidelity for every layer and head.

Jane: It’s essentially a targeted denoising mechanism that corrects errors caused by losing information during compression, ensuring the retained data remains as accurate as possible for us to use in the downstream tasks.

Meng: If I were building this system, I'd be excited about how OBF provides a mathematically grounded way to ensure the communication channel integrity, even when we're aggressively trimming bandwidth in production. The math backs up the engineering choice.

Lu: I think this allows us to design highly complex reasoning chains where each agent can rely on a compact but perfectly informed snapshot of its predecessor’s work, knowing that information is preserved structurally.

Lalam: It represents an architectural improvement that ensures our AI systems are more reliable and less prone to the pitfalls of information loss in the long-term deployment of agents.

Paper discussion segment 3: Tom: That leads us into a deeper look at "When Less Latent Leads to Better Relay," focusing on how the mechanics of Orthogonal BackFill (OBF) are specifically designed to solve this challenge.

Jane: The authors aren't just compressing; they are intentionally creating a targeted denoising effect by isolating the parts of the deleted information that were orthogonal—meaning they were completely separate—from what was retained.

Lu: This is a beautiful conceptual leap, suggesting that our AI models have internal states so rich that we can mathematically separate "useful" signal from noise even when we force them to lose data.

Meng: The engineering benefit of this is the precision: OBF doesn's just dump random data back in; it uses attention weights to decide exactly *where* and *how much* correction is needed, which makes deployment much more predictable.

Lalam: This specificity ensures that our AI systems are not only efficient but also trustworthy, making it a huge step toward building agents whose interactions are reliable over time.

Tom: The paper shows that by using this targeted backfill, the error bound is significantly reduced compared to just letting the data go, offering a clear path to maximizing fidelity in each individual step of every agent's process.

Jane: It’s a method of precision correction; instead of just hoping that compression didn't hurt us, we are actively fixing it by ensuring the retained data remains as accurate as possible for us to use in the downstream tasks.

Meng: I find the dynamic scaling aspect of OBF particularly clever, which is an important design feature for making sure the correction is always proportional to how much attention was originally paid to those deleted tokens.

Lu: This allows us to envision highly complex reasoning chains where each agent can rely on a compact but perfectly informed snapshot of its predecessor’s work, knowing that information is preserved structurally.

Lalam: It represents an architectural improvement that ensures our AI systems are more reliable and less prone to the pitfalls of information loss in the long-term deployment of agents.

Conclusion: Tom: We've spent so much time breaking down this research today that it's hard to imagine where we go from here without summarizing what we've seen in "When Less Latent Leads to Better Relay: Information-Preserving Compression for Latent Multi-Agent LLM Collaboration."

Jane: The biggest realization is that the need for massive data exchange between agents isn't a fundamental law of AI; we have found a way to be smarter about how we process and use the information that is already there.

Tom: That’s exactly it, Jane; this research shows us how high performance and low communication cost are not mutually exclusive goals in this latent multi-agent setting.

Meng: The practical implications for building large-scale systems are enormous because this gives us a roadmap to drastically reduce the computational footprint of complex AI agents in production environments.

Lu: I find it incredibly exciting because it suggests our agent designs can move from merely being data pipelines to becoming truly intelligent, highly selective information processors across different domains.

Jane: I agree with Lu; the paper proves that quality of matters more than quantity when we are dealing with these multi-agent interactions.

Tom: It is a major milestone in achieving this level of efficiency while maintaining peak performance for every single benchmark tested.

Meng: This means we can start designing systems where the bandwidth constraint is no longer the limiting factor, which is crucial for scaling up AI deployments.

Lu: I think it opens up new pathways for creating massive, complex reasoning chains that were previously limited by the sheer volume of communication we had to send between agents.

Lalam: It represents a powerful shift where our AI systems are becoming leaner and more reliable, requiring far less constant correction in our daily interactions with technology.

Tom: These findings have truly redefined what's possible in multi-agent design, and we're going to have to keep researching this topic for next week.

Yiping Li, Zhiyu An, Wan Du

University of California, Merced · Department of Computer Science and Engineering at University of California, Merced (as part of the affiliation)

cs.LG

Submitted: 2026-08-22

Updated: 2026-08-25

Importance score: 81/100

The gist: Problem Statement and Motivation In Large Language Model (LLM) based multi-agent systems (MAS), communication often moves beyond natural language to preserve internal computation, such as in

Key concepts

Information-Preserving Compression
These methods (L-OBF/H-OBF) compress data by ensuring that the resulting information remains highly accurate. This allows AI systems to leverage necessary context without requiring massive state transfers, proving that high performance and low communication cost are not mutually exclusive goals.
Orthogonal BackFill (OBF)
OBF is the core mechanism for minimizing information loss during compression. It identifies unique components of deleted data that do not overlap with retained values, then injects them back into the model weights using a dynamic scaling factor for precise correction.
Latent Multi-Agent LLM Collaboration
This refers to advanced AI setups where multiple large language model agents work together on complex tasks. The primary challenge addressed is managing the sheer volume of communication required between these agents, aiming to make interactions reliable and efficient.

Terminology

Summary

Problem Statement and Motivation

In Large Language Model (LLM) based multi-agent systems (MAS), communication often moves beyond natural language to preserve internal computation, such as in LatentMAS (LMAS). In this setting, agents relay its full key-value (KV) cache directly to the next agent allowing downstream reasoning to continue from internal states. However, this high fidelity comes at a substantial cost, as relaying full KV caches consumes substantial memory and bandwidth.

The core challenge addressed in this work is that existing methods for reducing KV-cache cost—such as StreamingLLM and H2 O—target single-agent online cache management, not the relay problem in multi-agent latent communication. The authors propose a solution where compression occurs at the relay boundary between agents.

Methodology: Orthogonal BackFill (OBF)

The authors adapt KV-cache eviction methods to this inter-agent relay setting and introduce Orthogonal BackFill (OBF). The goal is that under a fixed relay budget, we aim to make each retained state carry as much signal as possible for the downstream agent.

1. Cache Decomposition

The multi-agent latent memory structure is formalized using a four-part decomposition:

  • Attention Sink (S): A persistent anchor inherited across the chain, that maintains numerical stability.

  • **Inherited Message History (H): The compressed historical context transmitted by upstream agents.

  • **Current Prompt Context (P = S Pi): The latent states induced by the task-specific instructions of the current agent.

  • **Current Latent Reasoning (G = SGi): The latent states generated by the current agent’s ongoing reasoning, which form the immediate causal basis for pi i 's generation.

The compression operator acts on these prompt states P, resulting in a retained subset P keep,i.

2. The Orthogonal BackFill Process

When hard eviction occurs (e.g via LMAS-Attn or LMAS-Attn), the deleted positions P del,i are identified. OBF addresses the information loss through a series of steps:

  • Construct Orthonormal Basis: An orthonormal basis for span(V keep) is constructed via a reduced QR decomposition on V keep:

Q = orth V keep

  • Calculate the Residual: The residual R is calculated by subtracting the projection of deleted values onto the retained span:

R = V del - (V del Q)Q T

  • Denoise and Summarize: The residual R is denoised using Singular Value Decomposition (SVD), retaining its top k right singular vectors (C). The deleted content is then summarized by an attention-weighted average:

= 1 over N sum j=1 N A t + epsilon w j R[j,:]

  • Project and Scale: The summary r is projected onto the residual principal subspace, and then scaled by the deleted-versus-retained demand ratio (delta):**

delta = del over keep + epsilon

  • Inject Correction: OBF modifies only the value states (V t*) by adding the same residual vector delta to all retained value states:

V t* from V t + delta, t in P keep

Why Orthogonality Matters: The authors emphasize that inject[ing] only the orthogonal component... avoids double counting of information already accessible through retained tokens. This ensures the correction is added to an otherwise empty subspace, preventing interference with the amplified retained signal.

Experimental Results and Performance

The method was evaluated across nine benchmarks spanning mathematical reasoning, expert and commonsense QA, and coding (GSM8K, AIME 2024/2025, GPQA, MedQA, ARC-E/C, MBPP-Plus, HumanEval-Plus).

  • Accuracy: With only 9.9%–20.2% of the prompt KV states retained, L-OBF and H-OBF achieved significant performance gains. Averaged across the nine benchmarks, L-OBF and H-OBF sit +1.9 and +1.7 pp above Full.

  • Efficiency: The compressed variants showed high inference efficiency (Table 2). While they produce more tokens than the full baseline, compressed relay produces about 10% more final-agent tokens for only about 5% more text-inference time.

Discussion and Conclusion

The results suggest that more information does not necessarily lead to better communication; preserving the most useful information matters more. The effectiveness of the compression depends on the task:

  • Benchmarks where Full relay leads (e.g, GSM8K, GPQA) reward factual retention.

  • Benchmarks where compressed variants lead (e.g., AIME 2024/2025, MBPP+) demand multi-step reasoning or code generation, where the prompt acts more like a scratchpad and redundant context distract[s].

The authors conclude that compressed relay is the rational default once communication cost matters, viewing OBF as a frontier-improving operator on average rather than a universal Pareto optimum.

Improvements for AI systems

As a diligent AI researcher, I have rigorously analyzed the provided paper to derive actionable, high-impact improvements for optimizing multi-agent LLM collaboration systems. The core of the paper is that structured compression and orthogonal backfilling are superior to simple hard eviction when replacing full KV cache relay.

The following improvements focus on transforming a standard Latent Multi-Agent System (LMAS) architecture into an optimized, cost-efficient system using the principles of Orthogonal BackFill (OBF).


The primary improvement is replacing the naive hard eviction step (Gen) with a sophisticated, signal-preserving compression operator (OBF). This requires three integrated components: Selection, Residual Calculation, and Dynamic Injection.

Instead of uniformly applying one compression strategy across all tasks, we implement a Task Selector Module that determines the optimal granularity based on the task type:

  • Headwise Selection (H): Used for tasks requiring specialized signal (e.g., mathematical reasoning, coding). This retains tokens that receive high attention from the current latent reasoning states (G), allowing specialized heads to retain their critical, focused information.

  • Layerwise Selection (L): Used for tasks requiring a holistic overview (e.g., complex QA/commonsense). This aggregates attention across heads within a layer, providing a smoother, more robust summary of the information retained across different functional components of the model.

For every agent's local prompt states (P i), we implement a dynamic backfilling process that goes beyond simple deletion:

  • Residual Isolation: For the deleted value states (V del (the portion of KV cache removed), we calculate the orthogonal residual R = (the deleted values) - (projection onto the retained span). This captures information that is genuinely new or unrepresented by any retained token.

  • Dynamic Demand Scaling: We calculate a scaling factor alpha = del / keep, where represents the attention mass of deleted vs. retained tokens. This ensures that the backfill is proportional to the actual demand for that specific information, preventing over-correction in low-attention contexts.

  • Injection: The final correction vector delta is injected uniformly into all retained value states (V keep to V keep + delta).

The system must manage the variable cache lengths and the state transition across agents efficiently:

  • Variable Cache Handling: The communication protocol must explicitly track and align the retained/deleted positions using a precise per-sample position cursor (c) to ensure that when one agent compresses its KV cache, the receiving agent can correctly map all subsequent tokens.

  • Position ID Consistency: We maintain explicit position IDs for both latent steps and text generation, ensuring that no real token in any downstream agent can mistakenly attend to a pad-origin KV slot, even if the compression operator changes the cache length between agents.

By implementing this refined architecture, the improved AI system will demonstrate significant advantages over standard full-KV relay systems:

  1. Substantially Reduced Communication Cost: The system achieves high accuracy while only retaining a small fraction (9.9%–20.2%) of the original prompt KV states, drastically reducing memory and bandwidth consumption compared to the full-KV baseline.

  2. Guaranteed Information Fidelity: Unlike simple eviction, the system guarantees that critical information is preserved or compensated for by injecting an orthogonal residual (delta). This minimizes hard-eviction error (the gap between the original message and the compressed output).

  3. Task-Adaptive Performance: The system can dynamically choose between Layerwise and Headwise compression, allowing it to excel in specific domains:

  • Mathematical/Coding Tasks: Benefit from Headwise selection, retaining specialized signal necessary for complex logic.

  • Complex QA Tasks: Benefit from Layerwise selection, providing a robust summary of the entire context.

  1. Optimized Accuracy-Efficiency Frontier: The system achieves a superior balance between accuracy and cost, consistently outperforming the full-KV baseline in critical benchmarks (e.g., AIME24, MBPP+), demonstrating that less information can indeed lead to better communication when the retained information is maximally useful.

Sources

Related papers