ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping
cs.AI
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/Czzzk/Resispec
Terminology
Sources
- GPT-4 Technical Report
- Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding
- Adaptive Input Representations for Neural Language Modeling
- Controlling Computation versus Quality for Neural Sequence Models
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- Accelerating Large Language Model Decoding with Speculative Sampling
- Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
- Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
- Speculative Decoding and Beyond: An In-Depth Survey of Techniques
- Towards Optimal Multi-draft Speculative Decoding
- SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
- EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
- Language Models are Few-Shot Learners
- Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
- Instruction Tuning with GPT-4
- Efficiently Scaling Transformer Inference
- Accelerating LLM Inference with Staged Speculative Decoding
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection