Entropy-Aware Token Rejection for Improving Speculative Decoding
cs.CL, cs.AI
Submitted: 2025-12-29
Updated: 2026-08-30
Terminology
Sources
- GPT-4 Technical Report
- Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding
- Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment
- Qwen Technical Report
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- Accelerating Large Language Model Decoding with Speculative Sampling
- EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
- A Survey on Collaborative Mechanisms Between Large and Small Language Models
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Measuring Mathematical Problem Solving With the MATH Dataset
- Universal Model Routing for Efficient LLM Inference
- RewardBench: Evaluating Reward Models for Language Modeling
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
- EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- Reward-Guided Speculative Decoding for Efficient LLM Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering