FaST: Feature-aware Sampling and Tuning for Personalized Preference Alignment with Limited Data
cs.CL
Submitted: 2025-08-06
Updated: 2026-09-01
Comments: EMNLP 2025 - Main Conference
Code: https://github.com/facebookresearch/ELI5
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: LLM-powered conversational assistants are often deployed in a one-size-fits-all manner, which fails to accommodate individual user preferences.
Terminology
Abstract
LLM-powered conversational assistants are often deployed in a one-size-fits-all manner, which fails to accommodate individual user preferences. Recently, LLM personalization -- tailoring models to align with specific user preferences -- has gained increasing attention as a way to bridge this gap. In this work, we specifically focus on a practical yet challenging setting where only a small set of preference annotations can be collected per user -- a problem we define as Personalized Preference Alignment with Limited Data (PPALLI). To support research in this area, we introduce two datasets -- DnD and ELIP -- and benchmark a variety of alignment techniques on them. We further propose FaST, a highly parameter-efficient approach that leverages high-level features automatically discovered from the data, achieving the best overall performance.
Sources
- A General Language Assistant as a Laboratory for Alignment
- PERSONA: A Reproducible Testbed for Pluralistic Alignment
- Less is More: Improving LLM Alignment via Preference Data Selection
- A Survey on Personalized Alignment -- The Missing Piece for Large Language Models in Real-World Applications
- Direct Language Model Alignment from Online AI Feedback
- Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging
- Data-Efficient Alignment of Large Language Models with Human Feedback Through Natural Language
- Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
- Drift: Decoding-time Personalized Alignments with Implicit User Preferences
- The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
- Personalized Language Modeling from Personalized Human Feedback
- Proximal Policy Optimization Algorithms
- FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
- ALMA: Alignment with Minimal Annotation
- SPRI: Aligning Large Language Models with Context-Situated Principles
- Orchestrating LLMs with Different Personalizations
- PersonalLLM: Tailoring LLMs to Individual Preferences
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering