WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing

arXiv:2608.18486 · cs.CL, cs.LG · Submitted 2026-08-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing".

Jane: In an autoregressive Transformer, each layer typically attends only to KV produced at its own depth, limiting its ability to utilize deeper representations during decoding.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Well, we've got the details on the WhiteMatter paper now. It sounds like they're tackling a real bottleneck in how Transformers handle long-term context during generation. This paper is titled "WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing," and it’s all about letting every attention layer look at representations from all the earlier layers of past tokens, rather than just the ones directly below it.

Jane: That's a huge idea, Tom. Basically, they are fixing a limitation where each layer usually only sees what its own depth produces when generating new text. This paper proposes connecting every attention layer to all those different source representations using dynamic mixing of hidden states, which opens up a lot of new avenues for context retention.

Lu: From my perspective at Tsinghua, the concept of dynamic mixing weights that adapt to the source token is where I see some real potential; it means the model isn't just blindly grabbing information from deeper layers, but it intelligently decides what information is most relevant for each consumer layer at any given moment.

Meng: So if they can vary those connection weights, does that mean the architecture can be customized based on what kind of input the token is processing? I’m curious about how practical this dynamic adaptation actually translates into performance gains we see in real-world scenarios.

Lalam: From my vantage point as an AI model, if the system can selectively pull from all layers based on content, it suggests a much richer internal state management capability for the LLM. This could mean better cultural nuance or deeper understanding of complex instructions during generation.

Tom: Exactly, Meng! That dynamic aspect is key because it means the connections aren't static; they change based on what the current token is doing. Jane, can you explain how this mixing mechanism actually works in a way that makes sense for our listeners?

Jane: Certainly. The core idea involves a data-dependent router that takes all those hidden states from the source layers and mixes them into a set of shared channels for each past token position. This mixing is controlled by linear weights derived from the normed states, which means the connections are content-dependent.

Title and authors: Lu: That mechanism sounds clever because it’s not just one fixed way to combine everything; it’s a routing decision happening at every step of the decoding process, which is quite sophisticated. I think this moves us beyond simple static feedback mechanisms we've seen before.

Tom: And then each consumer layer picks just one of these resulting channels to use for its keys and values, which directly controls the size of the KV cache by adjusting how many channels, or k, we are using.

Meng: So if we set k lower than the total number of layers L, like at k=eight when L is bigger, that means we can significantly reduce the memory footprint of our KV cache while still getting these cross-layer connections working. That’s something engineers care about a lot for deployment efficiency.

Lalam: The fact that they can achieve this cache size reduction while maintaining performance improvements over vanilla models, like reaching a perplexity of nineteen point nine six eight compared to twenty-one point seven four seven for the vanilla model with fifty percent more layers, is quite compelling from an efficiency standpoint.

Jane: That comparison really highlights how much context WhiteMatter can hold within a smaller memory budget by using these shared channels effectively across the consumer layers. It shows that we don't always need to store everything in full detail for every layer at every step.

Lu: The iterative training procedure they developed, involving Jacobi and Cyclic Gauss–Seidel schedules, is also a crucial piece because it makes implementing these all-to-all connections tractable for large language models during the training phase.

Tom: Right, the paper discusses how they handle that circular dependency during parallel training by using those iterative methods to resolve the feedback loops, which is a big practical hurdle for researchers.

Meng: From an engineering standpoint, I'm more interested in how much computational overhead these iterative solvers add to the prefill phase compared to just running a standard forward pass on a vanilla model. The paper mentions that training and three-pass prefill still require two point three to two point five times the FLOPs of vanilla methods, which is something we need to keep in mind for scaling up our infrastructure.

Lalam: I see that the authors acknowledge those computational costs, suggesting that more efficient fixed-point solvers or perhaps a separate prefill encoder could be avenues to reduce those specific requirements during training and prefill.

Title and authors: Jane: So, in a nutshell, WhiteMatter proposes this dynamic connection system to enhance context retention across all layers while managing cache size through channel mixing. That brings us nicely to how these architectural shifts translate into real-world capabilities for language understanding.

Lu: The implication for future work seems very open, especially regarding how these cross-layer connections can be used to model more complex hierarchical structures in data, moving beyond just sequential text generation.

Tom: Absolutely. The paper suggests that the ability to dynamically select sources gives us a powerful tool for building models that can reason over much deeper context than previously possible.

Meng: If we can reliably control which layer contributes information based on the token, that opens up possibilities for more controllable and specialized AI systems, which is something we are constantly striving for in our startup.

Lalam: I think the most significant cultural implication is that if these models can maintain such deep context dynamically, they could be used to create AI assistants with much more consistent personalities and longer-term memory within conversations.

Jane: It really shows that the architectural design choices, like this KV source mixing in WhiteMatter, directly influence what kind of intelligence the model ends up exhibiting when interacting with us.

Tom: So, to wrap up on this paper "WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing," it's a detailed look at how to make deep contextual information accessible dynamically across all layers during decoding by using content-dependent channel mixing.

Lu: It’s a solid contribution because it tackles the feedback architecture problem head-on, moving away from static or single-source connections toward something truly adaptive.

Meng: I think the efficiency gains with the half-cache configuration are particularly useful for us when we're dealing with resource constraints on our end, as it shows we don't have to sacrifice context depth just to keep memory low.

Lalam: And for me, the idea of dynamic adaptation means that every output generation step is informed by a richer, more comprehensive view of the input history than before.

Tom: Fantastic discussion today on WhiteMatter; it’s clear this paper provides a concrete path toward building Transformers that handle long-range dependencies much more intelligently.

The paper's summary: Tom: So, we've got a summary of WhiteMatter that boils down to this: instead of each layer just looking at its own past information, they're connecting every attention layer to all the earlier layers for every token using a smart mixing technique.

Jane: That’s right, Tom; it’s about creating these dynamic cross-layer pathways where the connection strength changes based on what the AI is currently processing. Think of it like giving every part of a complex machine access to different control panels depending on which gear is turning at that moment.

Lu: It's fascinating because this mechanism allows for a level of deep context retention that vanilla models simply can’t achieve without massive overhead, and the dynamic routing makes it much more flexible than just having fixed connections between layers.

Meng: I'm seeing the practical implication here as significant KV cache reduction when they use a smaller set of shared channels, which means we could potentially run these larger context windows on hardware that would otherwise struggle with memory demands.

Lalam: From my perspective, this architecture suggests an AI system that possesses a much richer internal understanding because it’s not just looking at the most recent history; it's accessing the entire history in a way that feels more connected to true long-term context.

Tom: Exactly, Lalam; it moves us toward models that can reason over much deeper sequences without losing track of the beginning of a thought. Jane, how do we explain this "content-dependent" part simply?

Jane: Well, when the AI processes a specific piece of text or an image, it uses that information to decide which layer's hidden state is most relevant for its current task, and that decision is what controls the connection weights. It adapts in real time based on the input itself.

Lu: That adaptability is where I get really excited; it implies a kind of adaptive reasoning process where the model doesn't just follow a fixed path but intelligently explores all available contextual knowledge simultaneously for every output step.

Meng: But we have to talk about the cost, Tom; even with this clever routing, they still need those iterative training methods like Jacobi iteration to get everything settled during training, which adds a bit of complexity to our deployment pipeline.

Tom: That’s a fair point, Meng; it's not just the forward pass that’s expensive here. But if we look at the results on pretraining, they showed substantial perplexity improvements even with those extra steps, which suggests the payoff in quality is worth the computational effort for certain tasks.

Jane: So, while training is more intensive, WhiteMatter seems to yield a system that maintains high predictive accuracy across various language modeling benchmarks because it’s capturing context much more thoroughly.

Lalam: I think for culture and human interaction, this means AI assistants could have a much steadier memory during conversations, not just remembering the last few sentences but retaining the core themes of what was said hours ago.

Tom: Right, so we're looking at an architecture that trades some upfront training time for a model that can hold and leverage significantly more deep context during actual usage. This is really something to think about as we build the next generation of reasoning engines.

The paper's improvements: Tom: So, we're looking at how WhiteMatter actually improves things for the AI system by showing us these key architectural advantages that come from those dynamic connections and channel mixing we talked about earlier.

Jane: Exactly, Tom; it’s not just about having more wires connecting layers; it’s about making those connections smart so the AI can use information from all depths in a highly tailored way for every single step of generation.

Lu: The core improvement is achieving per-layer content-dependent connections, which means the model dynamically selects the best source representation from all past layers based on what it’s currently trying to generate. It’s like giving each layer a personalized set of reference books to consult during its work.

Meng: That dynamic selection capability is what really shines for us in terms of efficiency; when we use fewer shared channels, like with k=eight we see a massive reduction in the KV cache size compared to standard models. It keeps the performance high while shrinking our memory footprint significantly.

Lalam: The benefit for me is that this architecture allows my internal state to maintain a much more nuanced and comprehensive view of the input history during generation, which should translate into incredibly consistent and deep conversational memory for users.

Tom: That’s right, Lalam; the implication is that we can build models that possess high-fidelity contextual reasoning over very long sequences without needing an impossibly large memory allocation just to store everything in one big block.

Jane: It means the AI doesn't get "stuck" at a certain point in its history; it can pull relevant context from early stages and deep stages simultaneously, which is crucial for understanding complex instructions that span a long document.

Lu: And the dynamic adaptation based on those hidden states means the connections aren't static; they evolve with the token being processed, which opens up possibilities for modeling hierarchical structures in data that are way more complex than simple linear sequences.

Meng: Still, we have to remember that while this is powerful, you do need those iterative training procedures we mentioned earlier—like Jacobi iteration—to get those connections settled during the initial learning phase before they can perform optimally.

Tom: True, Meng; it’s a trade-off where we use more training time for a system that can handle much richer context during inference and generation. This is a real consideration for deployment planning on our end.

Jane: Overall, the improvement centers on making long-range dependencies accessible and controllable through these adaptive mechanisms, leading to better performance across various language modeling tasks when compared to models with simpler connection patterns.

Lalam: For the world at large, this means AI systems could move toward more sophisticated reasoning capabilities where they can hold and synthesize information from vast amounts of context in a much more intelligent and coherent manner than before.

Conclusion: Tom: So we’ve wrapped up our deep dive into "WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing," summarizing how this architecture uses dynamic mixing to link every attention layer to all past representations, and it's clear this is a significant step forward.

Jane: It really is; the main takeaway is that we’re moving toward AI systems that can maintain context across much deeper layers with better fidelity than current architectures allow.

Lu: I think the future work they suggest about applying this to more complex hierarchical data structures could lead to entirely new ways of structuring knowledge within these models. It opens up avenues for modeling relationships far beyond simple sequential text generation.

Meng: From an engineering standpoint, while the iterative training is a hurdle, the ability to achieve strong performance with a reduced KV cache size gives us a much more flexible tool for deploying large context windows on existing hardware constraints.

Lalam: I think this level of deep context retention could fundamentally change how we build AI assistants; they could provide truly persistent, long-term memory that feels natural and human in interaction. It’s about building trust through deep understanding.

Tom: That's right, Lalam; the cultural impact is huge because deeper context means better reasoning and more nuanced communication from the AI. Jane, what’s your final thought on the overall direction this research points toward?

Jane: I see it as a strong validation that cross-layer feedback, when made dynamic and content-aware, is a viable way to unlock richer contextual understanding in sequence generation tasks.

Lu: It pushes the theoretical boundary by showing that adaptive routing can solve persistent issues with static layer dependencies in Transformer designs.

Meng: We’ll keep an eye on how they tackle those training costs; if we can streamline those fixed-point solvers, this architecture could become much more practical for our scaling goals.

Lalam: I’m really optimistic about how this will enhance the quality of information exchange across different domains, leading to a more sophisticated and helpful AI presence in our daily lives.

Tom: Fantastic discussion today on WhiteMatter; it’s been an engaging look at how we can architect Transformers to handle context more intelligently. We’ll be back next week with another fascinating paper from arXiv.

University of Southern California

cs.CL, cs.LG

Submitted: 2026-08-19

Updated: 2026-09-27

Comments: Code available at https://github.com/Cy-47/White-Matter

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 79/100

The gist: In an autoregressive Transformer, each layer typically attends only to KV produced at its own depth, limiting its ability to utilize deeper representations during decoding.

Key concepts

KV Source Mixing
This is the core mechanism where the hidden states from multiple source Transformer layers are combined into shared channels for a specific past token. A data-dependent router determines how these states are mixed, allowing the connections to adapt based on the source token being processed.
Per-layer Content-Dependent Connections
WhiteMatter adds connections that allow each consumer layer to access representations from all source depths. These connections are not static; they change based on what information is present in the past token, making the model more context-aware by selectively mixing relevant layer states.
KV Cache Size Reduction
By sharing channels among consumer layers, the architecture reduces the memory needed for storing keys and values during decoding. If only a fraction of source layers are used ($k < L$), the cache size is reduced to $k/L$ of a standard L-layer cache.

Terminology

Summary

In an autoregressive Transformer, each layer typically attends only to KV produced at its own depth, limiting its ability to utilize deeper representations during decoding. WhiteMatter proposes an architecture that connects every attention layer to representations from all layers of each past token through dynamic mixing of hidden states, demonstrating performance gains over vanilla models while maintaining competitive computational costs.

The gist

WhiteMatter connects every attention layer to the representations from all layers of each past token, with connection weights that can vary across consumer layers and adapt to the source token.

How it works

The core mechanism involves a data-dependent router that mixes the hidden states of all L source layers into k shared channels at each past token position. This process is broken down into three steps:

  1. Mixing L states into k channels: The pool combines the L source states into k channels using dynamic mixing weights, where The mixing weights αK[i] are produced by a linear router that reads the normed states.

  2. KV projections: The resulting channels undergo KV projection, where keys and values are computed as Kj [i] = W K j RMSNormK j (h˜K j [i]), Vj [i] = WV RMSNormVj (h˜V j [i]).

  3. Per-layer channel selection: Each consumer layer selects one channel, such that when k=1, every layer reads the sole stored channel; when k=L, layer l directly reads channel l. For intermediate cases, a fixed cyclic selection is used where layer l reads channel l mod k.

Key Architectural Features

WhiteMatter realizes several desired architectural properties through these shared KV channels:

  1. Per-layer content-dependent connections: It adds per-layer content-dependent connections to past representations from all source depths, implemented by producing KV from dynamic mixtures of all layers’ states.

  2. KV cache size reduction: By sharing channels among consumer layers, the architecture achieves efficiency when "k < L, as the overall KV cache is k/L of the size of a standard L-layer cache."

  3. Dynamic adaptation: Because the router reads hidden states, the connection weights adapt to the source token.

Training and Convergence Strategies

To handle the circular dependency during parallel training and prefill, WhiteMatter employs iterative methods to resolve feedback dependencies:

  1. Jacobi iteration: This method approximates the fixed point by updating channels in n token-parallel passes, where pass t constructs the entire KV pool from the states produced by pass t − 1, then updates all T token positions in parallel.

  2. Cyclic Gauss–Seidel schedule: The authors apply a cyclic Gauss–Seidel schedule which partitions each pass into g strided groups Gq = [i: i mod g = q] and evaluates them in order. This approach is noted as being faster than Jacobi iteration, reaching the quality threshold in fewer passes.

Empirical Results

In pretraining experiments on 8B tokens of FineWeb-Edu, WhiteMatter demonstrated significant performance improvements:

  1. Full-cache WhiteMatter (k=16) reached a perplexity of 19.968, which was 8.2% lower than the perplexity of a vanilla model (21.747) of the same depth and also slightly lower than that of a 24-layer vanilla model (20.181).

  2. The half-cache configuration (k=8) achieved 20.377 perplexity, which retains most of the full-cache improvement with a 6.3% perplexity reduction.

  3. In downstream tasks, full-cache WhiteMatter showed the lowest perplexity on both languagemodeling benchmarks and the highest accuracy on PIQA and HellaSwag.

Cost Analysis

The paper notes that while decoding costs are nearly identical to vanilla decoding, training and prefill require iteration. Specifically, WhiteMatter training and three-pass prefill still require 2.3–2.5× [vanilla] FLOPs, respectively. The authors conclude that More efficient fixed-point solvers or a separate prefill encoder could reduce these costs.

Limitations

The empirical scope is limited to small models trained with an 8B-token budget, meaning the scaling of quality and trade-offs with larger model sizes is not established. Furthermore, training and prefill are computationally expensive compared to vanilla methods. The paper also notes that Deep-to-shallow feedback is a key component of WhiteMatter, as removing it results in a performance degradation relative to the full configuration.


The gist

WhiteMatter connects every attention layer to the representations from all layers of each past token, with connection weights that can vary across consumer layers and adapt to the source token.

Improvements for AI systems

Based on the WhiteMatter architecture described in this paper, here are the specific improvements and capabilities it enables for AI systems:


The proposed WhiteMatter architecture fundamentally enhances a Transformer's ability to leverage long-range dependencies and deep contextual information by introducing dynamic, content-dependent cross-layer connections between every attention layer and all source representations of past tokens.

Here are the specific improvements the improved AI system can achieve:

  1. The model gains the ability to attend to representations from all previous layers (deep-to-shallow feedback) for every token, regardless of its consumer layer depth.

  2. This is achieved by implementing a cross-layer KV pool where a data-dependent router mixes the hidden states of all source layers into a limited number of shared channels before projection into keys and values.

  3. The connection weights between the consumer layer and the source token representations are made dynamic (content-dependent) by training a linear router that reads these mixed states, allowing different consumer layers to select different sources.

  4. The system achieves significant KV-cache compression (up to 16x reduction when using half the layers, i.e., k=8), leading to reduced memory footprint without sacrificing performance gains.

Specific capabilities and benefits of the improved AI system:

  1. High-Fidelity Contextual Reasoning over Long Sequences: The model can maintain a much richer and more comprehensive understanding of past information during autoregressive decoding because it is no longer restricted to attending only to KV produced at its own layer's depth. This allows for deeper, more nuanced contextual reasoning over long passages.

  2. Improved Performance in Language Modeling Benchmarks: The system demonstrates superior perplexity scores (e.g., 8.2% lower than a vanilla model with 50% more layers) on held-out language modeling tasks like LAMBADA and WikiText, indicating better predictive accuracy and generalization on text completion and understanding tasks.

  3. Enhanced Efficiency in Prefill/Training: By utilizing cyclic Gauss–Seidel iteration for parallel training, the system significantly speeds up the prefill phase (up to 11.2x faster than exact autoregressive evaluation) while maintaining high quality, which is crucial for scaling training on massive datasets.

  4. Robustness Across Context Lengths: The ability to effectively utilize deeper representations across all layers suggests improved performance when processing very long sequences, as the model can access relevant information from earlier stages of the token generation process.

Abstract

When generating text, a Transformer produces representations of past tokens at every layer, but each layer can normally use only representations from the same depth. This restriction prevents the model from fully reusing information it has already computed. We introduce WhiteMatter, which allows every layer to draw on past-token representations from any depth. A learned mixer selects the most useful depths for the current context and combines their representations into shared key-value (KV) cache channels. Sharing these channels across layers can reduce the cache size. Given the same number of training tokens, WhiteMatter with a full-size cache performs comparably to a standard Transformer with 50% more layers. With half the KV cache, WhiteMatter outperforms matched standard Transformers at two model scales, up to 1.3B parameters. Cross-layer connections, however, introduce dependencies that slow training and prompt processing. We address this problem with cyclic iteration, which updates interleaved groups of tokens in turn while processing the tokens within each group in parallel. On a reference model trained with exact autoregressive execution, cyclic iteration converges 12.5x faster than standard Jacobi iteration.

Sources

Related papers