Hindsight-Anchored Policy Optimization: Learning Through Hindsight with Thompson Sampling-Inspired Adaptive Gating
cs.LG, cs.AI, cs.CL
Submitted: 2026-03-11
Updated: 2026-09-14
Comments: Published as a conference paper ICLR 2026 CAO Workshop
Code: https://github.com/huggingface/open-r1
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement Learning with Verifiable Rewards improves reasoning in large language models, yet on-policy learning often suffers from cold-start challenges in sparse-reward settings.
Terminology
Abstract
Reinforcement Learning with Verifiable Rewards improves reasoning in large language models, yet on-policy learning often suffers from cold-start challenges in sparse-reward settings. Recent mixed-policy approaches address this by combining off-policy teacher data with on-policy training. However, simply combining these introduce a persistent off-policy gradient mass that risks training collapse and instability. To address this challenge, we propose Hindsight-Anchored Policy Optimization (HAPO), a framework that allows teacher intervention to act as a temporary support. HAPO employs Beta-Binomial confidence gating, an adaptive gating mechanism that decides when to open the gate for teacher intervention. The intervention operates with Synthetic Success Injection, which replaces the group's lowest-reward rollout with a verified teacher trajectory. We also introduce adaptive threshold annealing, which gradually retracts the support and restores on-policy training within a finite horizon to mitigate persistent off-policy drift. We demonstrate that HAPO can be layered on top of existing mixed-policy methods in a generalizable manner. Across six math reasoning benchmarks and two model scales, HAPO improves the average accuracy of three major mixed-policy methods while maintaining training stability.
Sources
- Hindsight Experience Replay
- SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Measuring Mathematical Problem Solving With the MATH Dataset
- Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- UFT: Unifying Supervised and Reinforcement Fine-Tuning
- Understanding R1-Zero-Like Training: A Critical Perspective
- Towards a Unified View of Large Language Model Post-Training
- Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
- Reinforcement Learning: An Overview
- Training language models to follow instructions with human feedback
- Proximal Policy Optimization Algorithms
- Trust-Region Adaptive Policy Optimization
- Finetuned Language Models Are Zero-Shot Learners
- Learning to Reason under Off-Policy Guidance
- Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
- A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks