ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients
cs.LG, cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/euReKa025/ORPG
Terminology
Sources
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Adapting Auxiliary Losses Using Gradient Similarity
- Pareto Multi-Objective Alignment for Language Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
- Solving Quantitative Reasoning Problems with Language Models
- GD$^2$PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization
- DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
- Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
- Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
- O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
- Training language models to follow instructions with human feedback
- Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks