Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads

arXiv:2601.21351 · cs.LG, cs.AI · Submitted 2026-08-14 · Read on arXiv

Chendong Song, Meixuan Wang, Hang Zhou, Hong Liang, Yuan Lyu, Zixi Chen, Yuwei Fan, Zijie Zhou

Hong Kong University of Science and Technology · Huawei Hong Kong Research Center · Tsinghua University · Peking University

cs.LG, cs.AI

Submitted: 2026-08-14

Updated: 2026-08-18

Comments: Submitted to Neurips 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 68/100

Terminology

Summary

Summary

This paper addresses the problem of optimally provisioning resources in Attention–FFN disaggregated (AFD) LLM serving architectures under stochastic workloads. AFD is an emerging architecture that separates stateful, KV-cache–dominated Attention computation from stateless, compute-intensive FFN computation, connected by per-step communication. The central design question is determining the optimal ratio of Attention instances to FFN instances (denoted r in an rA–1F topology), since missizing induces step-level blocking and costly device idle time.

Problem and Motivation. The paper notes that existing systems set r through empirical search or naive deterministic approximations that ignore workload stochasticity. A principled analytical provisioning framework is still lacking. The difficulty arises because "the Attention workload is fundamentally stochastic. FFN computation depends only on batch size and remains stable across decode steps. In contrast, Attention workload evolves continuously—the KV cache grows with each step, and completed requests are replaced by new ones with variable prompt lengths."

Architecture Background. The paper describes the AFD execution mechanics: In an AF-disaggregated deployment, each Transformer layer comprises three distinct stages: Attention computation, FFN computation, and inter-device Communication (A ↔ F). Attention blocks are stateful—during each decode step, every continuing request generates one output token, and its corresponding key-value is appended to the KV cache—while FFN blocks are stateless—their computation depends only on the current activations, not on sequence history. Microbatch pipelining is used to hide communication latency, but Attention's stateful nature introduces dynamic variability that disrupts this ideal schedule, creating pipeline bubbles.

Mathematical Model. The paper considers a general xA–yF topology, defining r:= x/y, where r does not need to be an integer, for example, r = 3.5 corresponds to a 7A–2F configuration. Each Attention instance maintains a microbatch of B requests. The per-step cycle time is modeled as τ(B; r) = max αA W B,r + βA, t C(rB), t F(rB), where W B,r is the barrier load (the maximum token load across the r Attention workers). The optimization objective is throughput per instance: Throughput per-inst(B; r) = (1/(r+1)) · (rB/τ̄(B; r)), where τ̄ is the expected cycle time.

Key Theoretical Contributions.

(i) Stationary per-slot token load (Lemma 4.1). Using a discrete-time renewal-reward theorem applied to one decode slot, the paper derives the stationary token load moments. The mean is θ = E[DP + D(D−1)/2]/E[D], and the second moment is E[Y2] = E[DP2 + PD(D−1) + D(D−1)(2D−1)/6]/E[D]. For independent P and D, θ = μ P + (μ D − 1)/2 + σ2 D/(2μ D). The paper emphasizes that it is not the arrival-average μ P + μ D (a natural but incorrect first guess), but the stationary age-adjusted load, with decode-length variance entering through length-biasing.

(ii) Barrier-aware Attention load (Theorem 4.3). The paper proves that (W B,r − Bθ)/(√B ν) ⇒ M r as B → ∞, where M r is the maximum of r independent standard normals, and E[W B,r] = Bθ + √B ν κ r + o(√B), with κ r = E[M r] ≈ √(2 log r) for large r. This quantifies the synchronization overhead from cross-worker stragglers.

(iii) Mean-field provisioning rule (Theorem 4.4). The closed-form optimal A/F ratio is obtained by evaluating candidate ratios: r* mf ∈ min(√(β C/(α C B)), √(β F/(α F B))), (μ A − β C)/(α C B), (μ A − β F)/(α F B), (β C − β F)/(B(α F − α C)), choosing the one with the largest throughput.

(iv) Barrier-aware refinement. The Gaussian cycle time approximation is τ G(B; r) = G B,r + σ A ∫ z B,r∞ (m − z B,r) rφ(m)Φ(m) r−1 dm, and the refined optimal ratio is r* G ∈ arg max Thr G(B; r), evaluated via one-dimensional search.

Validation. The paper develops a trace-calibrated AFD simulator implementing a discrete-event, cycle-by-cycle simulation with a six-state finite state machine and two batches in flight. Using latency parameters calibrated for DeepSeek-V3 on Huawei Ascend 910C NPUs (α A = 0.00165 cycles/token, β A = 50 cycles, α F = 0.083 cycles/request, β F = 100 cycles, α C = 0.022 cycles/token, β C = 20 cycles), with B = 256, μ D = 500, μ P = 100, the theoretical optimal ratio r* ≈ 9.3 closely matches simulation results. The paper reports that the predicted optimal ratio matches the simulation-optimal within 10%. A residual gap of up to 15% between mean-field prediction and simulation at large r is attributed to synchronization overhead, consistent with the barrier-aware analysis showing 11% overhead at r = 24.

Ablation Studies. The paper examines the impact of batch size B (optimal r* = 7.08, 9.34, 10.31 for B ∈ 128, 256, 512) and workload distributions, finding that the optimal r* scales with total context length—longer prefills and decode sequences require more Attention instances.

Practical Recipe. The paper provides a three-step procedure: (i) estimate θ̂ and ν̂2 from request traces using nonparametric estimators; (ii) compute r* mf in closed form; (iii) refine via r* G when cross-worker imbalance is non-negligible. The framework requires no parametric distributional assumptions, only finite moments, and the paper notes that in practice, serving systems impose maximum context and generation lengths, restoring all moments.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems, and what the improved systems can do:

1. Implement a closed-form provisioning rule for Attention–FFN disaggregated (AFD) LLM serving

  • Use Theorem 4.4 to compute the optimal Attention-to-FFN ratio r* directly from hardware latency coefficients and workload statistics, without empirical search.

  • The rule decomposes into three operating regimes (Attention-bottleneck, communication-bottleneck, FFN-bottleneck), each with an interpretable balance condition.

2. Add a barrier-aware synchronization correction

  • Use Theorem 4.3 and equation (9) to quantify cross-worker straggler overhead via a Gaussian order-statistic model.

  • This corrects the mean-field throughput estimate by accounting for the slowest Attention worker, which is critical as the number of Attention instances r grows.

3. Calibrate workload statistics nonparametrically from request traces

  • Use the estimators in Appendix A.6 (equations 15–16) to compute and squared directly from serving logs, without assuming a parametric distribution for prefill or decode lengths.

  • This makes the provisioning rule robust to arbitrary, real-world request length distributions.

4. Incorporate geometric decode-lifetime specialization

  • When decode lengths are approximately geometric (as observed in production traces, Figure 5), use Corollary 4.5 to obtain closed-form expressions for theta and nu squared, reducing calibration to simple mean and variance computations.

5. Build a trace-calibrated AFD simulator for validation

  • Use the discrete-event simulator described in Section 5.1 to validate provisioning decisions before deployment, with two-batch interleaved execution to hide communication latency.

  • Automatically determine the optimal number of Attention instances per FFN instance (the A/F ratio) for a given hardware platform and workload, within 10% of the simulation-optimal value, using only a request trace and hardware latency coefficients.

  • Predict per-instance throughput, TPOT, and idle ratios as functions of the A/F ratio, enabling capacity planning and resource allocation before deployment.

  • Adapt provisioning dynamically as workload characteristics (prefill length, decode length, batch size) change, by re-estimating theta and nu squared from recent traces and recomputing r* in milliseconds.

  • Quantify synchronization overhead due to cross-worker stragglers, allowing operators to decide whether load-balancing routing policies (e.g., Chen et al., 2026) are worth implementing.

  • Serve LLMs with higher hardware utilization by balancing Attention and FFN workloads, reducing idle time on both sides, and avoiding pipeline bubbles caused by KV-cache growth.

  • Operate across hardware platforms by calibrating the six latency parameters (alpha A, beta A, alpha F, beta F, alpha C, beta C) from first-principles hardware specs (Appendix B) or from execution traces, making the framework portable to GPUs, NPUs, and other accelerators.

Abstract

Attentio-FFN disaggregation (AFD) is an emerging architecture for LLM decoding that separates state-heavy, KV-cache-dominated Attention computation from stateless, compute-intensive FFN computation, connected by per-step communication. While AFD enables independent scaling of memory and compute resources, its performance is highly sensitive to the Attention/FFN provisioning ratio: mis-sizing induces step-level blocking and costly device idle time. We develop an analytical provisioning framework for AFD bundles in an r A-- 1 F topology under stochastic workloads. Two sources of randomness shape the problem: per-slot Attention workload evolves as KV caches grow and completed requests are replenished with random prompt and decode lengths, and synchronized execution across Attention workers introduces a barrier governed by the slowest worker. We address both via a renewal-reward characterization of the per-slot stationary token load, identifying a single workload statistic theta that governs provisioning under arbitrary prefill-decode distributions and admits a nonparametric estimator from request traces. The analysis yields a closed-form mean-field rule for the optimal A/F ratio decomposing into Attention-, communication-, and FFN-bottleneck regimes, together with a Gaussian barrier-aware refinement that quantifies cross-worker synchronization overhead. A trace-calibrated AFD simulator supports the framework across workloads: the predicted optimal ratio matches the simulation-optimal within 10%. Together, these results provide a compact, calibratable account of how stochastic workload structure determines provisioning in disaggregated LLM serving.

Sources

Related papers