Learning to summarize user information for personalized reinforcement learning from human feedback
cs.LG, cs.AI
Submitted: 2025-07-17
Updated: 2026-08-26
Comments: ICLR 2026; 10 pages for main text
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: As everyday use cases of large language model (LLM) AI assistants have expanded, it is becoming increasingly important to personalize responses to align to different users' preferences and goals.
Terminology
Abstract
As everyday use cases of large language model (LLM) AI assistants have expanded, it is becoming increasingly important to personalize responses to align to different users' preferences and goals. While reinforcement learning from human feedback (RLHF) is effective at improving LLMs to be generally more helpful and fluent, it does not account for variability across users, as it models the entire user population with a single reward model, meaning it assumes that everyone's preferences are the same. We present a novel framework, Preference Learning Using Summarization (PLUS), that uses reinforcement learning (RL) to learn to produce text-based summaries of each user's preferences, characteristics, and past conversations. These summaries condition the reward model, enabling it to make personalized predictions about the types of responses valued by each user. Both the user-summarization model and reward model are trained simultaneously, creating an online co-adaptation loop. We show that in contrast to the standard Bradley-Terry model, summaries produced by PLUS capture diverse aspects of user preferences, achieving a 11-77/% improvement in reward model accuracy. Key strengths of PLUS are: (1) robust performance with new users and conversation topics, achieving a 25% improvement over the best personalized reward model technique used for RLHF; (2) zero-shot personalization with state-of-the-art proprietary models like GPT-4 (e.g., PLUS-summary-conditioned responses achieved a 72% win rate compared to 28% for default GPT-4o); (3) learning from flexible user contexts beyond preference labels, and (4) interpretable representation of users, enabling greater transparency and user control in pluralistic LLM alignment.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- vec2text with Round-Trip Translations
- UltraFeedback: Boosting Language Models with Scaled AI Feedback
- Is Independent Learning All You Need in the StarCraft Multi-Agent Challenge?
- Scaling Laws for Reward Model Overoptimization
- The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
- Personalized Language Modeling from Personalized Human Feedback
- A Shared Low-Rank Adaptation Approach to Personalized RLHF
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- Training language models to follow instructions with human feedback
- Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Proximal Policy Optimization Algorithms
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- The Trickle-down Impact of Reward (In-)consistency on RLHF
- A Roadmap to Pluralistic Alignment
- Learning to summarize from human feedback
- RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs
- The Surprising Effectiveness of PPO in Cooperative, Multi-Agent Games
- Benchmarking Large Language Models for News Summarization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks