Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation

arXiv:2608.12385 · cs.AI · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation".

Jane: The paper was written by Liming Liu, Mingze Wang and Tuo Zhao from Georgia Institute of Technology and Peking University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the channel, everyone. Today we’re digging into a paper that’s been making the rounds on arXiv — it’s called “Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation.” Jane, I’ve got to say, the title alone got me hooked. It’s talking about splitting how a model handles the prompt versus how it generates new tokens.

Jane: Right, Tom. And for our listeners who might not be deep in the weeds of language model internals — let me put it simply. When a model reads a long prompt, that’s called prefill. It processes everything at once, like reading a whole page. Then when it starts writing, one word at a time, that’s decode. These two jobs have totally different hardware demands. Prefill is like a math marathon — lots of parallel computation. Decode is more like waiting in line — it’s slow, sequential, and mostly limited by how fast memory can feed the processor.

Tom: And that mismatch is exactly what the authors are attacking. They’re saying, why do we have to use the same amount of computation for both phases? Why not give decode extra brainpower without slowing down the prompt reading? That’s the core idea here.

Jane: Exactly. And the clever part is how they do it. They keep one main “flow” — the primary path — that handles the prompt and builds the memory cache. Then they add a second, auxiliary flow that only kicks in during decode. It reads from the same memory but never writes to it. So you get extra computation where it matters, without bloating the prompt processing or the persistent state.

Tom: And the authors are from Georgia Tech and Peking University — Liming Liu, Mingze Wang, and Tuo Zhao. They’ve got a really elegant way of sharing almost all the heavy weights between the two flows, so the extra compute doesn’t mean a whole new model. It’s like adding a second engine to the same car body.

Jane: That’s a good way to think about it. And the results — they consistently show lower validation loss across different model sizes and training budgets. We’re going to get into the nitty-gritty of how they actually built this and tested it, but the big picture is: this could change how we think about serving costs for large models.

Tom: I’m excited to dig into the architecture details next. Stay tuned, because this gets really interesting when they start talking about mixture-of-experts and how you can tune prefill and decode budgets separately.

Summary: Tom: We’re back, still on “Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation.” Jane, let’s talk about what the paper actually does, because the summary is pretty dense.

Jane: It is, but the structure is actually clean. They start with a standard Transformer — that’s the primary flow. It processes the prompt, writes the key-value cache, and does all the normal autoregressive work. Then they add a second trajectory — the auxiliary flow — that shares almost all the weights but has its own embedding table and a few learned coupling vectors.

Tom: And the key trick is that the auxiliary flow never feeds back into the primary flow. It only reads from it. So during prefill, you can just skip the auxiliary entirely. It only gets evaluated at the final prompt position and then at every generated token during decode.

Jane: Right. So you’re getting roughly double the computation at decode time, but the prompt processing cost stays the same as a standard model. And the KV cache — that’s the memory that stores all the attention keys and values — stays exactly the same size. No extra persistent state.

Tom: They also do something neat with the output. Instead of just reading from one flow, they train a mixture of the two flows’ next-token distributions. The weight on the primary flow is always at least half, so the model stays primary-dominated, but the auxiliary flow gets a direct training signal and a direct vote at inference.

Jane: And that mixture is what they use for both training and validation, so the numbers are consistent. They ran experiments on NanoGPT with different token budgets, on dense LLaMA-style models at different scales, and on sparse mixture-of-experts models. In every case, Dual-Flow beat the standard Transformer at the same token budget.

Tom: The numbers are pretty convincing. For example, in the NanoGPT scaling experiments, they fit a power law to the validation loss and found the asymptote drops from about two point nine four for the standard Transformer to about two point nine zero for Dual-Flow. That’s a real improvement in the floor, not just a head start.

Jane: And they didn’t stop there. They also did a training-compute comparison — basically, is it better to spend your compute on more tokens for a single flow, or split it across two interacting flows? The result was that Dual-Flow wins even when you match the approximate training FLOPs. That’s a strong statement about the value of parallel computation inside the model.

Tom: I want to get into the mixture-of-experts part next, because that’s where they really show off the phase decoupling. But before we move on — Jane, anything you want to highlight about the component ablations?

Jane: Just that they were careful. They compared against PHD, which is the closest prior work, and they showed that their minimal Dual construction already matches or beats it. Then adding the coupling vectors and the mixture readout each gave a small, consistent boost. It’s good science — they isolate each piece.

Tom: Alright, let’s move to the MoE stuff. That’s where the prefill-decode tradeoff gets really concrete.

Improvements: Tom: Welcome back to the show. We’re still on “Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation.” Jane, the mixture-of-experts section is where I think this paper really shines.

Jane: Absolutely, Tom. So in a standard MoE model, every token activates a small number of experts — say, four out of a hundred. The router picks which ones. In Dual-Flow, you have two routers: one for the primary flow and one for the auxiliary flow. And here’s the kicker — you can set them to different numbers of experts.

Tom: Right. So the primary flow might activate four experts during prefill, but the auxiliary flow can activate eight or sixteen during decode. That means prompt processing stays cheap, but the model gets more brainpower exactly when it’s generating the next token.

Jane: And they tested this in a fixed-prefill regime. They held the primary at four experts and varied the auxiliary from four to sixteen. The validation loss dropped monotonically as the auxiliary count went up. So more decode-side experts directly translate to better quality, without touching the prompt cost.

Tom: Then they flipped it around — fixed decode budget. They kept the total number of expert activations per decode step constant, but reallocated between primary and auxiliary. So you could have one primary and seven auxiliary, or four and four, or seven and one. This changes the prefill cost while keeping decode the same.

Jane: And the results there are fascinating. The best point wasn’t at either extreme — it was around three-quarters of the budget on the primary flow. That makes sense, because the primary flow carries the main representation, but the auxiliary still contributes directly to prediction. If you starve the auxiliary, you lose that extra vote.

Tom: They also introduced something called router replay. The idea is that the auxiliary router computes its own logits but only gathers weights at the expert indices selected by the primary router. That way, both flows use the same set of experts, so you don’t have to load extra expert weights during decode. You just apply the same experts to both hidden states.

Jane: And that’s a big deal for memory bandwidth. During decode, you’re usually bottlenecked by how fast you can read weights and cache from memory. If both flows reference the same experts and the same KV cache, you can load those once and serve both computations. It’s like carpooling for memory traffic.

Tom: They showed that router replay retains most of the gain from fully independent routing, while keeping the referenced expert set unchanged. That’s a practical win for serving systems.

Jane: And the broader point is that this architecture gives you a three-way tradeoff — prefill cost, decode cost, and quality. A serving system can pick the operating point that fits its workload. If you’re prefill-heavy, you can shrink the primary fan-out. If you have decode headroom, you can spend it on auxiliary experts.

Tom: I love that this isn’t just a theoretical idea — they actually swept the space and showed the frontier. Meng, I know you’re the engineer on our team — what do you think about the practical side of this?

Meng: Honestly, the memory-traffic sharing is the part that gets me excited. If you can double the arithmetic per byte loaded, that’s a direct win on decode throughput. The implementation is non-trivial, but the payoff is real.

Tom: Great point. Let’s wrap up with our final thoughts and what this means for the field.

Conclusion: Tom: We’re at the end of our discussion on “Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation.” Jane, let’s pull it all together for our listeners.

Jane: Sure, Tom. The core contribution is an architecture that lets you add computation during decode without touching the prompt processing or the persistent KV cache. The primary flow handles the prompt and writes the cache; the auxiliary flow only reads from it and adds a second opinion at every generated token.

Tom: And they showed consistent quality gains across NanoGPT data scaling, dense LLaMA-style models, and sparse MoE configurations. The mixture-of-experts work is especially compelling because it turns prefill and decode into independent knobs you can tune for your serving workload.

Jane: Right. And the training-compute comparison is important too — it’s not just about token efficiency. Even when you match the approximate FLOPs, Dual-Flow comes out ahead. That suggests the parallel computation itself is valuable, not just the extra parameters.

Tom: The component ablations and the comparison against PHD show they were rigorous. They didn’t just throw a new architecture at the wall — they isolated each piece and showed it earns its keep.

Jane: And the extensions — three flows, auxiliary-specific weights — show this is a design space, not a one-off trick. You can scale the auxiliary side independently.

Tom: So what’s the big takeaway for the field? I think it’s that we don’t have to treat inference as a fixed cost. We can shape it — spend more where it helps, less where it doesn’t. That’s going to matter a lot as models get bigger and serving costs dominate.

Jane: Absolutely. And it opens the door for workload-specific deployment. A chatbot with long prompts and short replies can be tuned differently than a code assistant that generates long outputs.

Tom: Alright, we’ve covered a lot of ground today. Thanks to everyone who tuned in. We’ll be back with another paper soon — until then, keep thinking about how we can make these models not just smarter, but more efficient where it counts.

Jane: See you next time, everyone.

Liming Liu, Mingze Wang, Tuo Zhao

Georgia Institute of Technology · Peking University

cs.AI

Submitted: 2026-08-17

Updated: 2026-08-18

Comments: 18 pages

Code: https://github.com/KellerJordan/modded-nanogpt

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 63/100

The gist: The paper introduces the Dual-Flow Transformer, an architecture that decouples prefill and decode computation in large language models.

Key concepts

Prefill
This is the initial phase where a model processes a long input prompt. It involves parallel computation, similar to reading a whole page, and is necessary for building the memory cache.
Decode
This is the generation phase where the model creates new tokens one by one. It is sequential and limited by memory bandwidth, making it suitable for adding extra computational resources.
Dual-Flow Transformers
The core architecture that handles both phases. It uses a primary path to process the prompt and build cache, and an auxiliary path that only reads from this cache during decode to add extra computation.
Mixture-of-Experts (MoE)
A technique where a model has multiple specialized components (experts). In Dual-Flow, both the prefill and decode phases can activate different numbers of these experts independently.

Terminology

Summary

The paper introduces the Dual-Flow Transformer, an architecture that decouples prefill and decode computation in large language models. The motivation is that the two inference phases place different demands on hardware: prompt prefill runs in parallel and tends to be compute-bound, whereas autoregressive decode is sequential and is often memory-bandwidth-bound. Conventional scaling raises both costs together, since every added layer is evaluated in both phases. The authors instead ask whether additional learned computation can be allocated to continuation prediction while preserving the prompt-wide primary computation and a single persistent key–value (KV) cache.

The architecture is described as follows: "Its primary flow is a complete causal language model that alone processes the prompt and writes the KV cache; the auxiliary flow can therefore be omitted across the prompt and activated only from its final position onward, where it adds continuation-prediction computation without writing persistent state or influencing the primary flow. The flows share all major attention, MLP, and output matrices and use separate token embeddings with lightweight coupling."

The design satisfies four properties simultaneously: "(i) a complete primary causal path that alone processes the prompt; (ii) no persistent auxiliary state, so the KV cache is unchanged; (iii) additional learned computation applied at continuation-prediction steps; and (iv) substantial sharing of layer weights and the primary cached state between the primary computation and the additional computation."

Technical construction. The two streams begin from distinct embedding tables, with the primary and auxiliary states denoted S1 and S2. All subsequent dense projections are shared. The primary attention is exactly causal self-attention. The auxiliary flow produces only a new query and uses a learned coupling vector al added to the primary query: Q̃2 = Q2 + al ⊙ Q1. Both attention computations use ordinary causal semantics and require no interleaved or method-specific attention mask. The MLP similarly uses coupling vectors bl and cl: H̃2 = H2 + bl ⊙ H1 and M2 = H̃2 WD + cl ⊙ M1. The interaction remains asymmetric: it enriches the auxiliary representation without introducing any dependence of the primary flow on the auxiliary flow.

The training objective is a mixture likelihood: Lt = −log(αt pt(yt) + (1 − αt)qt(yt)) where p and q are the primary and auxiliary next-token distributions. The mixture weight is αt = clip[0.5,1] Σv pt(v)2, which is a confidence score based on the primary distribution's concentration (its collision probability), clipped so the mixture weight on the primary flow is always at least one half.

For MoE models, the paper introduces router replay: Let the primary router produce logits r1 and top-k expert indices I1. The auxiliary router evaluates its own logits r2 but gathers routing weights only at the primary indices. This preserves a flow-specific routing signal rather than copying the primary probabilities, while keeping the set of referenced expert weights unchanged across flows.

Inference semantics. The serving procedure has three steps: (1) Prompt-wide prefill where the primary flow first processes positions 1,..., S exactly as a one-flow Transformer and writes their keys and values to the cache; (2) Continuation boundary where decoding begins with one auxiliary evaluation at the final prompt token; and (3) Autoregressive continuation where the primary flow appends one set of keys and values, both flows form queries over the shared cache, and the mixture predicts xt+1. The paper notes that the prompt-wide primary graph, FLOPs, and KV construction are identical to the one-flow Transformer, and the boundary step is equivalent to adding one token-level decode evaluation, rather than adding computation to prompt-wide prefill.

Training cost. "Every token position serves as the boundary for predicting its successor, so both flows are evaluated across the full training sequence... giving Dual-Flow approximately twice the dominant training FLOPs of the primary model. The only added non-embedding parameters are the per-layer coupling vectors, which account for less than 0.1% of the non-embedding parameter count. The second token-embedding table increases parameter and optimizer-state memory but contributes only an additional input lookup."

Resource accounting. "Because the two flows share weights and KV, grouping their states lets each shared matrix and primary KV region serve both computations. This increases the arithmetic performed per loaded weight and cached state without enlarging the unique data referenced by the decode step."

Experiments. The paper evaluates along three axes:

  1. NanoGPT token scaling. Using the modded-NanoGPT optimization setting with 12 layers, width 768, head dimension 128, and vocabulary 50,257, trained on FineWeb with a data multiplier D ∈ 1, 2, 3, 4, 5, 10, 20 setting the number of optimizer steps to 3,800D. Dual-Flow finishes below the standard Transformer at every budget, and the separation persists as the token budget grows. Fitting the scaling law l(D) = L + AD−gamma, the fitted asymptote decreases from 2.9416 for the Transformer to 2.9013 for Dual-Flow.

  2. Approximate training-compute match. Comparing the D = 20 Transformer with the D = 10 Dual-Flow run (since one Dual-Flow step uses approximately twice the shared-backbone training FLOPs), Dual-Flow nevertheless reaches lower validation loss, supporting the benefit of allocating training computation across two interacting flows.

  3. Component ablations and PHD comparison. The paper compares against PHD-2, the prior architecture closest to Dual-Flow in form. A minimal Dual variant (no a/b/c coupling, auxiliary-only readout) matches PHD-2 at smaller budgets and outperforms it as the data budget increases. Adding a/b/c coupling produces a small but consistent improvement, and restoring the probability mixture yields the complete model and the best performance across the tested budgets.

  4. Dense LLaMA model-size scaling. At scales 0.12B, 0.25B, and 0.5B parameters with training data set to 80 times parameter count, Figure 5 shows consistent Dual-Flow gains over the corresponding dense baselines throughout this scaling trajectory.

  5. Sparse LLaMA MoE and router replay. Using a Qwen-style sparse MoE, "Dual-MoE improves over standard MoE across the tested base configurations, both with router replay and with independent auxiliary routing. Replay retains most of the independent-routing gain while preserving the referenced expert-weight set."

Phase-specific MoE allocation. The paper studies two allocation regimes. In the fixed-prefill regime, k1 is fixed and k2 varies: validation loss decreases monotonically as the auxiliary expert budget grows. Thus, when prompt computation is fixed, additional continuation-side expert computation translates directly into better predictive quality. In the fixed-decode regime, k1 + k2 = K is held constant: "The half-prefill allocation improves over the corresponding standard MoE baseline in both slices. Reducing the prefill fraction to one quarter preserves competitive quality at the smaller budget and improves over the baseline at the larger one. The paper notes that the optimum can therefore lie inside the allocation interval rather than at either endpoint."

Extensions. The paper also explores another auxiliary flow and dense weights specific to the auxiliary flow. Three flows show lower validation loss than two-flow Dual-Flow at all five NanoGPT budgets. Auxiliary-specific dense weights (separate query and attention-output projections and separate MLP) lowers validation loss relative to shared-weight Dual-Flow throughout the measured NanoGPT range.

Contributions. The paper's contributions are: (1) the Dual-Flow architecture with asymmetric shared-KV attention, shared dense weights, separate embeddings, lightweight cross-flow coupling, and a mixture objective; router replay extends it to MoE; (2) empirical quality improvements across controlled data scaling, dense model-size scaling, and sparse base configurations; and (3) phase-specific MoE allocation where the primary and auxiliary expert fan-outs are independently configurable, exposing a trade-off among prefill cost, decode cost, and predictive quality.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:

  1. Phase-decoupled architecture: Implement a Dual-Flow Transformer where a primary flow processes the entire prompt and writes the KV cache, while an auxiliary flow is activated only at continuation-prediction steps. This decouples prefill (compute-bound) and decode (memory-bandwidth-bound) computation.

  2. Asymmetric shared-KV attention: The auxiliary flow generates only queries, not keys/values, and attends to the primary flow's cached keys and values. This doubles decode arithmetic without enlarging the unique data read per step.

  3. Weight sharing with lightweight coupling: Share all attention, MLP, and output matrices between flows. Add per-layer learned coupling vectors (al, bl, cl) that transfer information from primary to auxiliary at query, MLP-intermediate, and MLP-output levels. This adds <0.1% non-embedding parameters.

  4. Mixture next-token objective: Train with a mixture likelihood L = -log(α·p(y) + (1-α)·q(y)), where α is a confidence score based on the primary distribution's collision probability, clipped to [0.5, 1]. This ensures the primary flow remains dominant while the auxiliary contributes directly to prediction.

  5. Router replay for MoE: For sparse Mixture-of-Experts models, have the auxiliary flow reuse the primary flow's selected expert indices but compute its own routing weights at those indices. This preserves the referenced expert-weight set while applying experts to both flows' states.

  6. Phase-specific expert allocation: Make primary and auxiliary expert fan-outs (k1 and k2) independently configurable. This enables workload-specific trade-offs: fix prefill compute and increase decode experts, or fix decode compute and reallocate expert budget between flows.

  7. Separate token embeddings: Use independent embedding tables for each flow (E1, E2), which are the only additional vocabulary-scale parameters. This allows different representational starting points without adding matrix-based computation.

For inference efficiency:

  • Process long prompts (e.g., 100k+ tokens) with identical prefill cost to a standard Transformer, since the auxiliary flow is omitted during prompt processing.

  • During decoding, perform 2× the arithmetic per generated token without reading additional weights or KV cache entries, improving arithmetic intensity and potentially reducing memory-bound stalls.

  • Serve prefill-heavy workloads (e.g., batch processing, document analysis) with no auxiliary overhead, while still gaining quality from auxiliary computation at generation time.

For model quality:

  • Achieve lower validation loss than standard Transformers at matched token budgets. Specifically:

  • NanoGPT: 0.04 lower final validation loss at D=5 (3.068 vs 3.108).

  • Dense LLaMA: consistent gains from 0.12B to 0.5B parameters.

  • MoE: 0.05-0.10 lower validation loss with router replay.

  • Extend to 3+ flows for further gains (e.g., 3 flows beat 2 flows by 0.02 at D=5).

  • Optionally use auxiliary-specific dense weights for additional quality improvement (0.01-0.02 lower loss).

For deployment flexibility:

  • In prefill-constrained scenarios (e.g., long-context agents), fix k1=4 and increase k2 from 4 to 16, reducing validation loss by 0.03-0.05 without increasing prompt cost.

  • In decode-constrained scenarios (e.g., high-throughput serving), fix k1+k2=8 and reduce k1 to 2 (quarter prefill expert arithmetic) while maintaining or improving quality over standard top-8 MoE.

  • Choose optimal prefill fraction (e.g., k1/K ≈ 0.75 for best quality at fixed decode budget) based on workload priorities.

For training:

  • Use the same architecture for training, with both flows evaluated at every position (2× training FLOPs), but achieve better loss-per-compute than spending that compute on additional tokens for a single flow (validated at D=10 vs D=20 comparison).

  • Maintain the same optimizer settings and schedules as baseline models, requiring no special training infrastructure.

Concretely, a deployed system could:

  • Use a 0.5B-parameter Dual-MoE with k1=4, k2=8 for interactive chat: prefill cost matches a standard top-4 MoE, but decode quality improves by 0.03 validation loss.

  • Use a 0.25B-parameter Dual-Flow with k1=2, k2=6 for batch document processing: prefill expert arithmetic is halved, decode quality matches standard top-8 MoE, and total throughput increases.

  • Scale to 3 flows with shared coupling for high-quality generation tasks (e.g., code completion, long-form writing), gaining 0.02 additional loss reduction over 2-flow Dual-Flow.

Abstract

As large language models serve more requests, cumulative inference cost is becoming increasingly important relative to one-time training cost. The two inference phases stress hardware differently: prompt prefill is parallel and typically compute-bound, whereas autoregressive decode is sequential and often memory-bandwidth-bound. Conventional width or depth scaling increases both costs together because every added layer is evaluated in both phases. We ask whether additional learned computation can instead be allocated to continuation prediction while preserving the prompt-wide primary computation and a single persistent key-value (KV) cache. We introduce the Dual-Flow Transformer. Its primary flow is a complete causal language model that processes the prompt and writes the KV cache. The auxiliary flow is omitted during prompt processing and activated only from the final prompt position onward, adding continuation-prediction computation without writing persistent state or influencing the primary flow. The two flows share major attention, MLP, and output matrices, while using separate token embeddings and lightweight coupling. Sharing weights and the primary cache also creates opportunities to reuse loaded weights and cached keys and values during grouped execution. Across matched-token comparisons, Dual-Flow achieves lower validation loss across architectures and data configurations. In MoE models, the separation makes primary and auxiliary expert fan-outs independent controls over prompt cost, continuation cost, and predictive quality. We study two regimes: increasing decode computation at fixed prefill expert computation, and reallocating a fixed decode expert budget between the two flows. These experiments expose a prefill-decode-quality trade-off and demonstrate the potential of phase-specific expert allocation.

Sources

Related papers