Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient Reasoning
cs.LG
Submitted: 2026-04-02
Updated: 2026-09-07
Comments: 20 pages, 4 tables, 6 figures, appendix included
Code: https://github.com/ServiceNow/PipelineRL
License: http://creativecommons.org/licenses/by/4.0/
The gist: Building general-purpose reasoning models using reinforcement learning with verifiable rewards (RLVR) across diverse domains has been widely adopted by frontier open-weight models.
Terminology
Abstract
Building general-purpose reasoning models using reinforcement learning with verifiable rewards (RLVR) across diverse domains has been widely adopted by frontier open-weight models. However, their training recipes and domain mixtures are often not disclosed. Joint optimization across domains poses significant challenges: domains vary widely in rollout length, problem difficulty, and sample efficiency. Further, models with long chain-of-thought traces increase inference cost and latency, making efficiency critical for practical deployment. We present Apriel-Reasoner, trained with a reproducible multi-domain RL post-training recipe on Apriel-Base, a 15B-parameter open-weight LLM, across five domains using public datasets: mathematics, code generation, instruction following, logical puzzles, and function calling. We introduce adaptive domain sampling that preserves target completed-rollout ratios despite heterogeneous rollout dynamics, and a difficulty-aware length penalty that, at no additional training overhead, encourages longer reasoning for difficult problems and shorter traces for easy ones. Trained with a strict 16K-token output budget, Apriel-Reasoner remains effective at a 32K output budget and improves over Apriel-Base on AIME 2025, GPQA, MMLU-Pro, and LiveCodeBench while producing 30-50% shorter reasoning traces. Among the evaluated open-weight models of similar scale, it improves the accuracy-token tradeoff using a fully public-data recipe.
Sources
- Phi-4-reasoning Technical Report
- Pixtral 12B
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- GLM-5: from Vibe Coding to Agentic Engineering
- The Art of Scaling Reinforcement Learning Compute for LLMs
- Kimi K2: Open Agentic Intelligence
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- TACO: Topics in Algorithmic COde generation dataset
- SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
- DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
- Olmo 3
- PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence Generation
- INTELLECT-3: Technical Report
- Qwen3-Omni Technical Report
- LLMs Can Learn to Reason Via Off-Policy RL
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
- BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks