From Rollouts to Recipes: Self-Contained Post-Training for LLMs
cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: 14 pages, 5 figures. Accepted at EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Post-training large language models usually applies a single training recipe to all samples, even though the model's own rollouts reveal different sample-level learning states.
Terminology
Abstract
Post-training large language models usually applies a single training recipe to all samples, even though the model's own rollouts reveal different sample-level learning states. We propose Self-Routing, a behavior-conditioned post-training framework that uses rollout correctness and confidence to decide how each sample should be optimized. Depending on its behavior state, a sample is routed to GRPO, on-policy self-distillation, regularization, or skipping, allowing training to adapt without external teachers, extra annotations, or additional sampling. Experiments on mathematical reasoning across Qwen3 and Qwen3.5 backbones show that Self-Routing consistently improves over uniform GRPO, uniform OPSD, fixed mixtures, and simpler routing baselines. Further analyses show that the routing distribution changes over training and reduces unnecessary updates on low-signal or already stable samples.
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Universal Self-Consistency for Large Language Model Generation
- Training Verifiers to Solve Math Word Problems
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
- Self-Specialized Teachers for Domain Post-Training
- Measuring Mathematical Problem Solving With the MATH Dataset
- RLTF: Reinforcement Learning from Unit Test Feedback
- Training Compute-Optimal Large Language Models
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- Reinforcement Learning via Self-Distillation
- VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models
- Scaling Laws for Neural Language Models
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Qwen3.5-Omni Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering