MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices
Eunjeong Kim, Yeong Jun Jeon, Myeonggyun Han
Kyungpook National University
cs.OS, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: Published in LCTES 2026
Journal ref: Proc. ACM LCTES 2026, 180-192 (2026)
Code: https://github.com/tatsu-lab/stanford_alpaca
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices Summary This paper presents MemSpec, a prediction-guided, memory-aware runtime system designed to
Terminology
Summary
MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices
Summary
This paper presents MemSpec, a prediction-guided, memory-aware runtime system designed to improve adaptive speculative decoding for large language models (LLMs) on memory-constrained edge devices. The authors identify a fundamental limitation in existing adaptive speculative decoding methods: the mismatch between draft selection and draft availability under tight memory budgets. While adaptive methods can improve token acceptance rates by dynamically selecting among multiple draft models, on edge devices, switching to a non-resident draft incurs substantial loading overhead that often negates throughput gains.
The paper makes three key observations motivating the work. First, static draft selection is insufficient due to variability in draft effectiveness across workloads and generation stages. Experiments show that Oracle-Static (selecting the best draft per prompt via offline evaluation) improves normalized acceptance by 40.3% on average over General-Static (using a single general-purpose draft), while Oracle-Dynamic (switching drafts within a generation) provides an additional 25.7% improvement over Oracle-Static. Second, on memory-constrained edge devices, switching cost fundamentally limits adaptive draft selection—loading a non-resident 400M-parameter draft model takes 2.7× longer than a single speculative decoding iteration. Third, exploration-based adaptation (e.g., multi-armed bandit methods) fails to improve throughput under memory constraints because frequent model loading dominates execution time, with MAB-Async spending 46.4% of time on average executing suboptimal drafts while waiting for preferred drafts to become resident.
MemSpec addresses these challenges through three integrated components: a Prediction Engine, a Draft Model Cache Manager, and a Runtime Controller. The Prediction Engine uses a fine-tuned BERT encoder to rank candidate drafts based on both the input prompt and recent generated tokens, avoiding costly online exploration. The Draft Model Cache Manager maintains a small resident working set of draft models under memory constraints, using a top-K policy with asynchronous prefetching and eviction that protects the active draft. The Runtime Controller orchestrates non-blocking decoding, always executing the best currently resident draft while preparing better candidates in the background.
The key design principle is decoupling draft selection from execution. MemSpec performs scheduling every N iterations (default N=4), amortizing prediction overhead while enabling overlap between decoding and asynchronous model loading. The runtime derives a target working set from predicted rankings but only selects the active draft from the currently resident set, ensuring non-blocking execution. This transforms draft adaptation into a staged, overlapped process rather than a sequence of stall-heavy immediate switches.
Experiments on a Jetson Orin Nano platform with GPTQ INT4 LLaMA-2 7B and Qwen2.5 7B target models, each with five 400M-0.5B parameter draft models (one general-purpose, four domain-specialized for code, math, law, and medical), demonstrate significant improvements. MemSpec improves steady-state generation throughput by 40.7% on average over state-of-the-art bandit-based adaptive methods (MAB-Async) and by 58.8% over static baselines, while achieving 95–97% of the Oracle-Dynamic upper bound. Execution breakdown analysis shows MemSpec reduces fallback execution (decoding with suboptimal resident drafts) to below 5.5% of runtime, compared to 49.4% for MAB-Async, with prediction overhead accounting for less than 3.9% of total runtime.
Component analysis reveals that prediction alone (Prediction-Only) improves throughput by 21.6% over General-Static but still falls short of full MemSpec by 30.5%, demonstrating that proactive cache management is essential—prediction quality does not directly translate into execution quality without residency management. Sensitivity analysis shows MemSpec performs best with moderate scheduling intervals (N=4 to 8), benefits more from longer output sequences (where intra-sequence variation provides more adaptation opportunity), and achieves most of its performance benefit at small cache capacities (K=2), with only modest improvement as K increases to 4 on a Jetson AGX Orin platform.
The predictor achieves 71.6% top-1 accuracy and 95.7% top-2 recall when using both prompt and recent tokens, confirming that MemSpec only requires sufficiently accurate ranking to maintain high-utility drafts in the resident set rather than perfect prediction. Including startup costs (prompt processing, prefill, initial loading), MemSpec still reduces end-to-end latency by 32.3% over General-Static and 24.9% over MAB-Async.
The paper concludes that the primary bottleneck in adaptive speculative decoding on edge devices is not draft selection quality but ensuring timely availability of effective drafts. By jointly optimizing draft selection and memory management, MemSpec captures most of the dynamic headroom of Oracle-Dynamic without requiring impractical offline enumeration, demonstrating that adaptive speculative decoding on edge devices is fundamentally a joint optimization problem over draft selection and memory management.
Improvements for AI systems
Improvements to AI Systems:
-
Memory-Aware Speculative Decoding Runtime: Integrate MemSpec's joint optimization of draft selection and cache management into LLM inference engines. The improved system can dynamically switch between multiple draft models (e.g., domain-specialized 400M-parameter models) on edge devices with limited VRAM, achieving 40.7% higher throughput than bandit-based methods by decoupling selection from execution and using asynchronous prefetching.
-
Prediction-Guided Draft Ranking without Online Exploration: Replace multi-armed bandit exploration with a fine-tuned BERT encoder that ranks candidate drafts based on prompt + recent tokens. The improved system can achieve 71.6% top-1 accuracy and 95.7% top-2 recall, avoiding the 46.4% wasted time on suboptimal drafts that MAB-Async incurs, while maintaining 95–97% of Oracle-Dynamic performance.
-
Non-Blocking Adaptive Scheduling with Staged Overlap: Implement a runtime controller that schedules draft switches every N iterations (default N=4) and always executes the best resident draft while loading better candidates in the background. The improved system can eliminate stall-heavy immediate switches, reducing fallback execution to below 5.5% of runtime (vs. 49.4% for MAB-Async) and keeping prediction overhead under 3.9%.
-
Top-K Resident Cache Management with Active-Draft Protection: Use a small resident working set (K=2) with eviction policies that protect the active draft. The improved system can maintain high-utility drafts in memory without thrashing, achieving most of the performance benefit at minimal cache capacity—critical for devices like Jetson Orin Nano with <8GB shared memory.
-
Context-Aware Draft Adaptation Across Generation Stages: Leverage the predictor's ability to use both input prompt and recent generated tokens to switch drafts mid-generation (e.g., from code to math draft as token distribution shifts). The improved system can capture the 25.7% additional gain of Oracle-Dynamic over Oracle-Static without offline enumeration, adapting to intra-sequence variation in draft effectiveness.
-
End-to-End Latency Reduction Including Startup: Optimize initial draft loading and prefill to overlap with prompt processing. The improved system can reduce total end-to-end latency by 32.3% over static baselines and 24.9% over MAB-Async, making speculative decoding practical for interactive edge applications with strict latency budgets.
Capabilities of the Improved AI System:
-
Runs LLMs (e.g., 7B parameter) on edge hardware with 5–10× smaller memory footprint than the target model, using a portfolio of small draft models.
-
Achieves near-oracle adaptive decoding performance (95–97% of Oracle-Dynamic) without knowing the best draft in advance.
-
Maintains steady throughput during generation even when the optimal draft changes (e.g., code→math→code), by proactively prefetching and evicting drafts.
-
Operates with <4% prediction overhead and <5.5% fallback execution, ensuring that most compute goes to accepted tokens.
-
Scales to different edge platforms (Jetson Orin Nano, AGX Orin) with tunable cache size (K) and scheduling interval (N), adapting to memory and compute constraints.
Abstract
Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to speculate multiple tokens, reducing expensive target model decoding steps. Its effectiveness depends heavily on draft selection, motivating adaptive methods that exploit variation across inputs and generation stages. On memory-constrained edge devices, however, these methods often fail to improve end-to-end throughput due to the overhead of switching between draft models. We identify a key limitation in this setting: the mismatch between draft selection and draft availability under tight memory budgets. To address this challenge, we present MemSpec, a prediction-guided, memory-aware runtime for adaptive speculative decoding on edge devices. MemSpec decouples draft selection from execution through proactive resident working-set management. A lightweight predictor estimates draft effectiveness from prompt and generation context, while a memory-aware scheduler reduces reactive model loading overhead. Experiments on a Jetson Orin Nano show that MemSpec improves steady-state generation throughput by 40.7% on average over state-of-the-art bandit-based adaptive methods while closely approaching the oracle upper bound.
Sources
- Program Synthesis with Large Language Models
- Accelerating Large Language Model Decoding with Speculative Sampling
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Qwen2.5 Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
Related papers
- Planarian: Managing Agent State with Statepoints
- VUDA: Enabling Controlled Spatial Sharing of Graphics and Compute on NVIDIA GPUs
- ActKV: Efficient LLM Agents through Action-Guided KV Cache Management
- GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference
- SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live