Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode
cs.AR, cs.LG, cs.PF
Submitted: 2026-08-24
Updated: 2026-08-24
Comments: 77 pages, 2 figures. Model weights and binaries: https://huggingface.co/tompoper/cflow
Code: https://github.com/ggerganov/llama.cpp
License: http://creativecommons.org/licenses/by/4.0/
The gist: Single-token autoregressive decode on CPUs is bound by memory bandwidth, not arithmetic: a modern CPU sustains roughly 1 TFLOP/s of compute but only about 50 GB/s from main memory, and each generated
Terminology
Abstract
Single-token autoregressive decode on CPUs is bound by memory bandwidth, not arithmetic: a modern CPU sustains roughly 1 TFLOP/s of compute but only about 50 GB/s from main memory, and each generated token must stream every active weight once. This report argues that the most effective response is to co-design the model architecture and the inference runtime together. It presents cflow, a CPU-first streaming engine, alongside a family of pipeline-native transformer architectures whose inter-layer dependency graphs are constructed to permit a vertical, stage-major execution schedule. cflow stores weights as L2-sized tiles in compute-consumption order, reads only the top-k experts of each mixture-of-experts layer, fuses projections, and executes a delay-aware schedule from per-model dependency parameters. Across five architectures trained on TinyStories, one (arch2 4 combined) achieves a 2.00x reduction in critical-path weight bandwidth (9.00 to 4.50 MB/token) within 0.24 perplexity of the best candidate, and the tile layout incurs 7.29x fewer L1-data read misses than a row-major baseline. On a 30.9-billion-parameter pipeline-native MoE, cflow decodes at 5.94 tokens/s (tok/s) on a 32-vCPU Ice Lake server, ahead of llama.cpp (4.75) and the vLLM CPU backend (1.65) on comparably sized dense models. Realizing the expert-delay window as asynchronous I/O overlap on a disk-resident expert tier yields a further net win of up to 1.68x, matching the overlap model within 1%. Measurement refutes one of the eight design claims and leaves a second inconclusive; both are reported in full, with the conditions under which they would hold.
Sources
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- LLM in a flash: Efficient Large Language Model Inference with Limited Memory
- ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware
- Accelerating Large Language Model Decoding with Speculative Sampling
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- Mixtral of Experts
- Fast Inference from Transformers via Speculative Decoding
- Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time
- OLMoE: Open Mixture-of-Experts Language Models
- GLU Variants Improve Transformer
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
- PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- FBNetV2: Differentiable Neural Architecture Search for Spatial and Channel Dimensions
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4