TIDE: Temporal Incremental Draft Engine for Self-Improving LLM Inference
cs.LG, cs.AI
Submitted: 2026-02-05
Updated: 2026-09-24
Code: https://github.com/sgl-project/specforge
Terminology
Sources
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
- BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
- Accelerating Large Language Model Decoding with Speculative Sampling
- Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
- HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Fast Inference from Transformers via Speculative Decoding
- Online Speculative Decoding
- PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
- SGLang: Efficient Execution of Structured Language Model Programs
- s1: Simple test-time scaling
- gpt-oss-120b & gpt-oss-20b Model Card
- Qwen3 Technical Report
- ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks