mzCache: On-Device LLM Memory Management under Multitasking
cs.OS, cs.DC, cs.LG
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: MobiCom 2026
Code: https://github.com/lz4/lz4
License: http://creativecommons.org/licenses/by/4.0/
The gist: On-device mobile Large Language Model (LLM) inference is gaining significant attention.
Terminology
Abstract
On-device mobile Large Language Model (LLM) inference is gaining significant attention. However, mobile devices operate in highly dynamic multitasking environments where users frequently switch between applications. This creates memory pressure, forcing LLM memory (model weights and KV cache) to be evicted by the operating system. When a new inference request arrives, the inference system must restore the evicted memory through slow storage reads or recompute the entire KV cache, severely degrading responsiveness. To address this, we present mzCache, an on-device LLM inference system with specialized memory management for multitasking environments. Under unpredictable memory pressure, mzCache elastically evicts LLM memory and leverages the unified memory of mobile SoCs to enable zero-wait inference on the GPU with concurrent CPU-side restoration. mzCache realizes this through restoration-oriented memory management: LLM memory is partitioned into fine-grained shared buffers to enable partial eviction and restoration with concurrent cross-processor access, while hybrid swap and backward-out eviction policies ensure low-latency restoration from any eviction state. Implemented on llama.cpp and deployed as an Android application, mzCache achieves 2.1-5.5 times reduction in Time-to-First-Token compared to storage-backed partial offload and demonstrates its effectiveness in real multitasking scenarios.
Sources
- Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
- EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes
- The Llama 3 Herd of Models
- Qwen3 Technical Report
- Gemma 2: Improving Open Language Models at a Practical Size
- PowerInfer-2: Fast Large Language Model Inference on a Smartphone
- ShareChat: A Dataset of Chatbot Conversations in the Wild
- LLM as a System Service on Mobile Devices
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
- A Review on Edge Large Language Models: Design, Execution, and Applications
Related papers
- MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices
- Planarian: Managing Agent State with Statepoints
- VUDA: Enabling Controlled Spatial Sharing of Graphics and Compute on NVIDIA GPUs
- ActKV: Efficient LLM Agents through Action-Guided KV Cache Management
- GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference
- SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading