LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference
cs.LG
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: 6 pages, 13 figures. To appear in the Proceedings of the 32nd Annual International Conference on Mobile Computing and Networking (MobiCom '26)
License: http://creativecommons.org/licenses/by/4.0/
The gist: On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM.
Terminology
Abstract
On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM. Prior systems exploit activation sparsity and offload weights to SSD or flash storage, but face a fundamental systems trade-off: accurate sparse execution decisions require the latest context, whereas efficient computation-I/O overlap requires early prediction. As a result, existing designs either serialize execution or incur redundant weight fetches, extra computation, and large cache overheads. We present LeanStream, a streaming speculate-and-refine framework for efficient on-device LLM inference. LeanStream progressively refines computation, loading, and cache-retention priorities using partial GPU results, enabling fine-grained overlap between GPU execution and storage I/O. We implement LeanStream on both mobile and embedded platforms. Compared with prior on-device LLM inference systems, LeanStream reduces memory usage by 4.8 times to 7.5 times at the best throughput achieved by prior work, while further improving token generation throughput by 1.6 times to 2.1 times.
Sources
- PowerInfer-2: Fast Large Language Model Inference on a Smartphone
- 6G Non-Terrestrial Networks Enabled Low-Altitude Economy: Opportunities and Challenges
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen3 Technical Report
- Learning Space Partitions for Nearest Neighbor Search
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Categorical Reparameterization with Gumbel-Softmax
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks