JuDi: Revisiting Judge Decoding from First Principles via Training-Free Distributional Divergence
cs.CL
Submitted: 2026-01-08
Updated: 2026-09-27
Terminology
Sources
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- Accelerating Large Language Model Decoding with Speculative Sampling
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
- Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
- The Llama 3 Herd of Models
- A Survey on LLM-as-a-Judge
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
- OpenAI o1 System Card
- Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
- Qwen3 Technical Report
- SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
- Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering