Correctness Forensics for Batch Speculative Decoding: Diagnosing the Ragged Tensor Problem
cs.CL, cs.AI
Submitted: 2025-10-26
Updated: 2026-08-31
Code: https://github.com/eBay/spec
Terminology
Sources
- LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
- REST: Retrieval-Based Speculative Decoding
- SPEED: Speculative Pipelined Execution for Efficient Decoding
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- Accelerating Large Language Model Decoding with Speculative Sampling
- Accelerating LLM Inference with Staged Speculative Decoding
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation
- TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
- Speculative Decoding: Performance or Illusion?
- The Synergy of Speculative Decoding and Batching in Serving Large Language Models
- Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
- Qwen3 Technical Report
- Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
- Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
- Speeding up Speculative Decoding via Sequential Approximate Verification
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering