Selective Off-Policy Reference Tuning with Plan Guidance
cs.AI
Submitted: 2026-05-12
Updated: 2026-09-24
Terminology
Sources
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
- A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
- Nudging the Boundaries of LLM Reasoning
- Kimi K2: Open Agentic Intelligence
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
- Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
- Self-Hinting Language Models Enhance Reinforcement Learning
- Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
- First Return, Entropy-Eliciting Explore
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
- Self-Distillation Enables Continual Learning
- RLPR: Extrapolating RLVR to General Domains without Verifiers
- S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
- Free Process Rewards without Process Labels
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection