SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference

arXiv:2608.13076 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Divya Jyoti Bajpai, Kishan Kumar Upadhyay, Manjesh Kumar Hanawal

Indian Institute of Technology Bombay

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: SPADE: Speculative Decoding for Precise and Low-Cost Distributed Edge–Cloud Inference Abstract Summary: The paper addresses the challenge of deploying Large Language Models (LLMs) constrained by

Terminology

Summary

SPADE: Speculative Decoding for Precise and Low-Cost Distributed Edge–Cloud Inference

Abstract Summary:

The paper addresses the challenge of deploying Large Language Models (LLMs) constrained by high computational demands. Deploying smaller LLMs directly on the edge degrades accuracy, while deploying larger cloud-based models preserves performance but incurs expensive per-token computation. The authors present SPADE, a distributed inference framework integrating speculative decoding (SD) across edge and cloud. A compact draft model on the edge generates candidate tokens rapidly, and a large verifier model on the cloud validates these tokens in parallel. Accepted tokens are retained, while only rejections trigger verifier correction, substantially reducing cloud queries. The plug-and-play design shifts computation to the edge, lowers inference time and cloud cost, and preserves accuracy without retraining. Experimental results on SpecBench and CNN/Dailymail datasets show SPADE reduces cloud model calls by 76% with zero loss in accuracy compared to the full model.

Introduction Summary:

LLMs achieve breakthroughs but demand vast computational resources rarely available on edge devices. Existing strategies like model pruning, weight quantization, and knowledge distillation reduce size but often reduce accuracy. Running full models on cloud servers restores performance but introduces cost that grows with resource utilization duration, as autoregressive generation requires sustained computation per token. The central challenge is reducing cloud reliance while maintaining full-model accuracy. Speculative decoding (SD) accelerates LLM inference while preserving accuracy using a two-model setup: a smaller draft model proposes tokens, and a larger verifier evaluates them in parallel. Verified tokens are retained, the first rejected token is replaced with verifier predictions, and generation continues. This cuts autoregressive structure—a 50-token sequence normally requiring 50 full model calls can often be completed with few calls. SPADE places the smaller model on the edge as draft generator and the full-scale model on the cloud as verifier, reducing cloud computational burden, inference latency, and expensive model calls while maintaining output fidelity without training.

Key Contributions:

  • Distributed Speculative Decoding Framework: A novel edge-cloud distributed inference setup leveraging speculative decoding to split computation between a smaller edge model and a full-scale cloud model.

  • Reduce Cloud Computation: Significantly reduces cloud model calls by using the edge model to generate draft sequences, lowering inference time and resource usage without compromising accuracy.

  • Maintain Full-Scale Accuracy: Ensures final outputs are equivalent to the full cloud model, guaranteeing fidelity with edge-side computation, without additional training.

  • Empirical evaluations: Results on SpecBench and CNN/DailyMail show cloud computation time reduced by 76% with no loss in performance.

Related Works Summary:

Distributed inference methods include layer-splitting (Neurosurgeon partitions DNNs by executing initial layers on edge and remaining on cloud), encoder-side training (head-network distillation), early-exit classifiers (SplitEE, I-SplitEE optimize splitting and prediction), and complexity-aware routing (DIMEE routes samples based on estimated complexity but relies on dataset-specific heuristics). Speculative decoding accelerates autoregressive models using a lightweight draft model and larger verifier for parallel validation; Self-Speculative Decoding (LayerSkip) reuses early layers of the large model as draft. Distributed inference methods face trade-offs: lenient routing reduces performance, strict routing increases cost, and complexity estimation lacks generalization. SPADE differs by applying speculative decoding to distributed inference with draft on edge and verifier on cloud, achieving zero performance loss relative to the large model (as proven in [12]), generalizing across tasks without dataset-specific heuristics, and being fully plug-and-play with no retraining.

Problem Setup Summary:

Autoregressive decoding in LLMs: The transformer architecture maps input tokens through embedding, L stacked transformer blocks, and a language modeling head. Each new token requires a full forward pass over the entire context, making inference latency dominated by model depth and memory requirements. Speculative decoding uses a two-model pipeline: (1) Drafting stage—the draft model autoregressively generates a block of d candidate tokens conditioned on the prefix; (2) Verification stage—the verifier evaluates the block in a single forward pass, accepting consistent tokens, replacing the first rejected token by sampling from an adjusted distribution, and resuming generation from the updated prefix. This ensures statistically identical output to verifier-only decoding while reducing expensive verifier calls.

SPADE Distributed Inference Setup:

  • Draft Model Mq on Edge: A compact LLM deployed on edge for fast inference under constrained memory and bandwidth, selected based on available resources (GPU memory, CPU throughput, bandwidth). Must be lightweight enough to generate speculative sequences without overwhelming the device.

  • Verifier Model Mp on Cloud: A larger, high-accuracy model that validates candidate tokens. Computational demands make it unsuitable for edge but ideal for cloud platforms that scale elastically. Minimizing verifier invocations directly reduces cost.

  • Verification Criterion: A candidate x ∼ q(x) is accepted with probability α(x) = min(1, p(x)/q(x)). If rejected, a replacement token is sampled from adjusted distribution p′(x) = norm(max(0, p(x) − q(x))), correcting draft bias and guaranteeing equivalence to verifier-only decoding.

  • Pipeline (Algorithm 1): Starting from a user prompt, the edge model generates a block of d draft tokens, sends them to cloud, where the verifier checks them in parallel in a single forward pass. Accepted tokens are appended; the first rejected token is replaced via sampling from p′(x); remaining draft tokens are discarded. Context updates with accepted and corrected tokens, and the draft model resumes generation until is generated. At least one token is appended per iteration, ensuring progress.

  • Draft token length (d): A key control parameter balancing computation and communication. Small d increases verification frequency and synchronization overhead; large d raises rejection likelihood and draft computation cost. d is selected empirically by monitoring acceptance rates on an initial validation subset (10 samples) to maximize throughput while maintaining stable acceptance rate.

Experiments Summary:

  • Datasets: CNN/DailyMail for summarization and Spec-Bench, which includes six subtasks (multi-turn conversation, summarization, translation, retrieval-augmented generation, question answering, mathematical reasoning).

  • Setup: Edge uses NVIDIA RTX 3080 GPU (12 GB RAM); cloud uses NVIDIA RTX A6000 GPU (48 GB RAM).

  • Models: LLaMA-3.2-1B as draft model at edge; LLaMA-3.1-8B as target verification model on cloud.

  • Metrics: For CNN/DailyMail, BLEU-1, BLEU-4, ROUGE-1, ROUGE-L, CIDEr-D (normalized to 0–100). For Spec-Bench, Gemini-2.5-Flash-Lite as automatic judge scoring responses on 1-5 Likert scale across six dimensions (correctness, instruction-following, completeness, clarity and coherence, conciseness, factuality).

  • Baselines: Target Model (large cloud model, upper bound in quality), Draft Model (lightweight edge model, low-latency outputs), and SPADE.

Results Summary:

  • Spec-Bench (Table I): SPADE achieves overall score 4.38 vs Target 4.45 and Draft 3.39. Task scores: Multi-turn Conversation 3.52 (Target 3.62, Draft 2.67), Translation 4.71 (Target 4.80, Draft 3.85), Summarization 4.55 (Target 4.61, Draft 4.41), Question Answering 4.15 (Target 4.25, Draft 3.08), Mathematical Reasoning 4.68 (Target 4.88, Draft 3.19), Retrieval-Augmented Generation 4.68 (Target 4.56, Draft 3.13). Efficiency: Mean Target Model Calls 30.16 (Target 133.25, Draft 0.00), Average Throughput 3.25 tokens/s (Target 2.43, Draft 3.91), Cloud runtime 0.23× (Target 1.00×). Cloud model calls reduced by 77.4%.

  • CNN/DailyMail (Table II): SPADE achieves BLEU-1 23.39 (Target 23.76, Draft 22.33), BLEU-4 06.98 (Target 07.57, Draft 06.49), ROUGE-1F1 37.99 (Target 38.38, Draft 36.05), ROUGE-LF1 23.92 (Target 24.32, Draft 22.49), CIDEr-D 03.19 (Target 02.50, Draft 01.15). Efficiency: Mean Target Model Calls 30.79 (Target 127.30, Draft 0.00), Average Throughput 1.95 tokens/s (Target 1.21, Draft 2.82), Cloud runtime 0.24× (Target 1.00×). Cloud calls reduced by 76%.

  • Analysis on d (Figure 2): Increasing d consistently reduces target model invocations as more draft tokens decrease verification frequency, amplified by strong alignment between draft and target models leading to higher acceptance rates. Poor alignment would reduce acceptance and increase target calls.

Conclusion Summary:

SPADE leverages speculative decoding in a dual-model setup with a lightweight draft model at the edge and larger verification model on the cloud. By using the verifier only for parallel verification rather than autoregressive generation, SPADE reduces costly cloud invocations. Experiments across multiple NLP tasks show SPADE significantly lowers cloud model calls while maintaining performance close to the full model. SPADE is established as a practical, scalable solution for latency-aware distributed inference, effectively balancing efficiency without any loss in performance.

Improvements for AI systems

Improvements to AI Systems:

  1. Cost-Aware Distributed Inference Engine: Implement a plug-and-play inference layer that automatically splits LLM workloads between edge devices (small draft model) and cloud (large verifier) using speculative decoding. The system reduces cloud API calls by 76% without retraining, cutting operational costs for LLM-powered applications (chatbots, summarizers, translators) while maintaining full-model output quality.

  2. Latency-Optimized Edge-Cloud Pipeline: Deploy a two-stage generation system where a 1B-parameter model on user hardware (e.g., RTX 3080) pre-generates token blocks, and an 8B+ model on cloud validates them in parallel. This reduces end-to-end inference latency by 4x (from 1.00x to 0.23x cloud runtime) compared to cloud-only inference, enabling real-time interactive applications on resource-constrained devices.

  3. Zero-Fidelity-Loss Model Serving: Use the acceptance-rejection sampling criterion (α(x) = min(1, p(x)/q(x))) to guarantee statistically identical outputs to the large cloud model, even when the edge draft model is much weaker. This allows enterprises to serve high-quality LLM responses without sacrificing accuracy, unlike quantization or distillation methods that degrade performance.

  4. Adaptive Draft-Length Controller: Integrate a dynamic token-block size selector (d) that monitors acceptance rates on a small validation set (e.g., 10 samples) and adjusts the draft length to maximize throughput. The system automatically balances communication overhead vs. rejection probability, making it robust to varying task types (e.g., math reasoning vs. summarization) and model alignments.

  5. Generalized Task-Agnostic Inference: Replace dataset-specific routing heuristics (e.g., complexity estimators) with a universal speculative decoding framework that works across six diverse NLP tasks (conversation, translation, RAG, QA, math, summarization) with no per-task tuning. This eliminates the need for custom split points or early-exit classifiers, making deployment trivial for new applications.

  6. Resource-Aware Edge Selection: Automatically choose the draft model size based on available edge hardware (GPU memory, CPU throughput, bandwidth) while ensuring the verifier model remains fixed on cloud. This enables deployment on everything from smartphones to workstations, with the system degrading gracefully—weaker edge models increase cloud calls but never break correctness.

  7. Throughput-Boosted Batch Processing: For non-interactive workloads (e.g., offline document summarization), the system increases token generation throughput by 34% (from 2.43 to 3.25 tokens/s) compared to cloud-only inference, by offloading speculative generation to edge GPUs. This accelerates bulk processing jobs while reducing cloud compute hours billed.

  8. Reliable Rejection Recovery: When the verifier rejects a draft token, the system samples from an adjusted distribution p′(x) = norm(max(0, p(x) − q(x))) to correct draft bias. This ensures the final output remains provably equivalent to the large model, preventing error accumulation over long generations (e.g., multi-turn conversations or long-form articles).

What the Improved AI System Can Do:

  • Serve LLM applications (chat, summarization, translation, RAG) on edge devices with cloud-level accuracy at 1/4th the cloud runtime and 1/4th the cloud API calls.

  • Operate without retraining or fine-tuning, working as a drop-in replacement for existing inference stacks.

  • Dynamically adapt to network bandwidth and device compute, maintaining stable throughput even under variable conditions.

  • Guarantee output fidelity to the large model, making it suitable for regulated domains (legal, medical) where accuracy is non-negotiable.

  • Scale to multi-user deployments by shifting compute burden to user devices, reducing central cloud infrastructure costs.

Abstract

Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands. Deploying smaller LLMs directly on the edge can circumvent this, but with degraded accuracy. Deploying smaller cloud-based big LLMs preserves performance, but at the cost of expensive per-token computation. We present a distributed inference framework,, that integrates speculative decoding (SD) across edge and cloud. A compact draft model deployed on the edge generates candidate tokens rapidly, and a large verifier model on the cloud validates these tokens in parallel. Accepted tokens are retained, while only rejections trigger verifier correction, substantially reducing the number of cloud queries. Our plug-and-play design shifts the bulk of computation to the edge, significantly lowers inference time and cloud cost, and preserves the accuracy of the big model without any retraining requirement. Our approach demonstrates a practical path toward scalable, cost-efficient, and accurate deployment of LLMs in real-world environments. Experimental results across multiple Natural Language Processing tasks using SpecBench and CNN/Dailymail datasets demonstrate that reduces the cloud model calls by 76% with zero loss in accuracy as compared to the full model.

Sources

Related papers