Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
cs.LG, cs.AI
Submitted: 2026-07-21
Updated: 2026-08-29
Comments: 24 Pages
Code: https://github.com/AgPriyank/OC-GRPO
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models.
Terminology
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives zero learning signal. Providing privileged guidance during training, such as solution prefixes, can help overcome this learning cliff by steering the model towards correct solutions with non-zero reward. We call these rollouts off-context: they are generated from a training prompt that contains privileged guidance, while the target objective is defined by the original prompt without that guidance. We introduce Off-Context GRPO (OC-GRPO), a minimally modified variant of GRPO that uses guided rollouts but applies an importance-corrected objective to steer the update back toward the original unguided objective, avoiding the mismatch that destabilizes uncorrected guided training. Empirically, our algorithm achieves a 3.8% absolute improvement (13.7% relative gain) over vanilla GRPO on average across standard mathematical reasoning benchmarks with negligible additional cost.
Sources
- RL for Reasoning by Adaptively Revealing Rationales
- Evaluating Large Language Models Trained on Code
- Self-Evolving Curriculum for LLM Reasoning
- Off-Policy Actor-Critic
- AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
- Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
- Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
- Reinforcement Learning via Self-Distillation
- OpenAI o1 System Card
- VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models
- Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
- POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
- Proximal Policy Optimization Algorithms
- Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks