GRPODropout: Less is More for Online Reinforcement Learning Rollouts
cs.LG, cs.AI, cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/hexuandeng/GRPODropout
Terminology
Sources
- Maximum a Posteriori Policy Optimisation
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities
- Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning
- Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR
- Language Models (Mostly) Know What They Know
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
- CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention
- Let's Verify Step by Step
- STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability
- Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index
- Training language models to follow instructions with human feedback
- Clip-Low Increases Entropy and Clip-High Decreases Entropy in Reinforcement Learning of Large Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Learning to summarize from human feedback
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks