CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning
cs.LG, cs.CL
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/HKUST-KnowComp/CorrGRPO
Terminology
Sources
- Program Synthesis with Large Language Models
- Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
- Evaluating Large Language Models Trained on Code
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Qwen2.5-Coder Technical Report
- MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
- Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- Training language models to follow instructions with human feedback
- ToolRL: Reward is All Tool Learning Needs
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Multi-objective Reinforcement learning from AI Feedback
- LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
- Qwen2.5 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks