Taming Speculative Search for Test-Time Scaling in LLM Serving
cs.DC, cs.CL, cs.OS
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/ihc-fan-lab/FastTTS
Terminology
Sources
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- $\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
- Training Verifiers to Solve Math Word Problems
- Distilling the Knowledge in a Neural Network
- Training Compute-Optimal Large Language Models
- Hydragen: High-Throughput LLM Inference with Shared Prefixes
- Scaling Laws for Neural Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- GPT-4 Technical Report
- GPT-4o System Card
- Solving math word problems with process- and outcome-based feedback
- OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing