To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals
cs.CL
Submitted: 2026-09-17
Updated: 2026-09-17
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Accelerating Large Language Model Decoding with Speculative Sampling
- CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
- Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
- The Llama 3 Herd of Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- Teaching Machines to Read and Comprehend
- BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- In-context Learning and Induction Heads
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models
- Qwen3 Technical Report
- DistillSpec: Improving Speculative Decoding via Knowledge Distillation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering