Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning
cs.AI
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/huggingface/trl
Project page: https://moonshotai.github.io/Kimi-K2
License: http://creativecommons.org/licenses/by/4.0/
The gist: Scaling test-time reasoning has substantially improved the problem-solving ability of large language models (LLMs), but standard autoregressive decoding still executes long reasoning traces
Terminology
Abstract
Scaling test-time reasoning has substantially improved the problem-solving ability of large language models (LLMs), but standard autoregressive decoding still executes long reasoning traces sequentially, creating severe latency for difficult tasks (up to days and weeks). Parallel reasoning offers a natural remedy. However, prior systems primarily focus on Subtask Parallelism, where the model learns to decompose a high-level task into smaller chunks that can be solved independently. This approach overlooks another pervasive form of parallelism: Trial Parallelism, where multiple speculative attempts explore, verify, and aggregate competing hypotheses in parallel. In this paper, we introduce Parason, which reveals and learns both forms of parallelism in LLM reasoning. Our analysis identifies Trial Parallelism as the majority of parallelizable reasoning computation (65.5% in DeepSeek-V4's reasoning steps in HLE), and it becomes increasingly dominant on hard problems. Guided by this taxonomy, Parason converts sequential reasoning traces into structured parallel trajectories with a context-free grammar, then trains models with Parallelism-Aware Group Relative Policy Optimization (PA-GRPO), whose reward jointly balances accuracy, latency, and the two parallelism ratios. At inference time, Parason executes the learned parallel structure through tool calls, translating theoretical savings to real-world wall-clock acceleration. Experiments on mathematical reasoning benchmarks including AIME24 and AIME25 show that Parason achieves an average acceleration about 1.7 times while maintaining competitive accuracy.
Sources
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
- Deep Think with Confidence
- Measuring Mathematical Problem Solving With the MATH Dataset
- PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
- Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
- Solving Quantitative Reasoning Problems with Language Models
- OpenAI o1 System Card
- Learning Adaptive Parallel Reasoning with Language Models
- Humanity's Last Exam
- HybridFlow: A Flexible and Efficient RLHF Framework
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- A Survey on Parallel Reasoning
- Qwen3 Technical Report
- Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge Generation
- SGLang: Efficient Execution of Structured Language Model Programs
- Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection