SEED: Self-Speculative Decoding via Implicit Encoder-Decoder
cs.CL, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/lhk2004/SEED
Terminology
Sources
- Speculative Streaming: Fast LLM Inference without Auxiliary Models
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- Accelerating Large Language Model Decoding with Speculative Sampling
- DFlash: Block Diffusion for Flash Speculative Decoding
- Training Verifiers to Solve Math Word Problems
- Parallel Token Prediction for Language Models
- Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
- Multi-Token Prediction via Self-Distillation
- DeepSeek-V3 Technical Report
- TiDAR: Think in Diffusion, Talk in Autoregression
- On multi-token prediction for efficient LLM inference
- Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
- Qwen3 Technical Report
- DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
- Self-Distillation for Multi-Token Prediction
- PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering