Recursive Self-Improvement via On-Policy Distillation for Reasoning
cs.CL
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Scaling Reinforcement Learning for Content Moderation with Large Language Models
- Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- OpenThoughts: Data Recipes for Reasoning Models
- Reinforced Self-Training (ReST) for Language Modeling
- The Geometry of Self-Verification in a Task-Specific Reasoning Model
- Understanding R1-Zero-Like Training: A Critical Perspective
- RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
- Qwen3 Technical Report
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Self-Distilled RLVR
- PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
- GRLO: Towards Generalizable Reinforcement Learning in Open-Ended Environments from Zero
- SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
- Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- From Emergence to Control: Probing and Modulating Self-Reflection in Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering