Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment
cs.LG, cs.AI, cs.CL
Submitted: 2026-02-17
Updated: 2026-08-31
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging
- gpt-oss-120b & gpt-oss-20b Model Card
- Personalized Language Modeling from Personalized Human Feedback
- Natural Emergent Misalignment from Reward Hacking in Production RL
- Evaluating ChatGPT as a Recommender System: A Rigorous Approach
- RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation
- Large Language Model Alignment: A Survey
- FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models
- Integrating Summarization and Retrieval for Enhanced Personalization via Large Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- Reinforcement Learning Enhanced LLMs: A Survey
- Qwen3 Technical Report
- Modelling and Analysis of Temporal Preference Drifts Using A Component-Based Factorised Latent Approach
- Group Sequence Policy Optimization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks