Planned Test-Time Scaling with Coordinated Reasoning Paths
cs.CL, cs.AI
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/shirley-wu/planned-test-time-scaling
Terminology
Sources
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models
- The Road Less Traveled: Enhancing Exploration in LLMs via Sequential Sampling
- Jointly Reinforcing Diversity and Quality in Language Model Generations
- Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
- Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
- On the Effect of Sampling Diversity in Scaling LLM Inference
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
- Qwen3 Technical Report
- Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering