Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
cs.LG, cs.AI
Submitted: 2025-04-27
Updated: 2026-09-22
Comments: Accepted at The 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026), Rabat, Morocco
Code: https://github.com/AAAhWei/Preference-Vector
License: http://creativecommons.org/licenses/by/4.0/
The gist: Ensuring that large language models (LLMs) are both helpful and harmless is a critical challenge, as overly strict constraints can lead to excessive refusals, while permissive models risk generating
Terminology
Abstract
Ensuring that large language models (LLMs) are both helpful and harmless is a critical challenge, as overly strict constraints can lead to excessive refusals, while permissive models risk generating harmful content. Existing approaches, such as reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO), attempt to balance these trade-offs but suffer from performance conflicts, limited controllability, and poor extendability. To address these issues, we propose Preference Vector, a novel framework inspired by task arithmetic. Instead of optimizing multiple preferences within a single objective, we train separate models on individual preferences, extract behavior shifts as preference vectors, and dynamically merge them at test time. This modular approach enables fine-grained, user-controllable preference adjustments and facilitates seamless integration of new preferences without retraining. Experiments show that our proposed Preference Vector framework improves helpfulness without excessive conservatism, allows smooth control over preference trade-offs, and supports scalable multi-preference alignment.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
- RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
- PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
- Mistral 7B
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Towards Safer Large Language Models through Machine Unlearning
- Enhancing LLM Safety via Constrained Direct Preference Optimization
- Benchmarking Batch Deep Reinforcement Learning Algorithms
- Proximal Policy Optimization Algorithms
- Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Steering Language Models With Activation Engineering
- Reinforcement Learning for LLM Post-Training: A Survey
- Ethical and social risks of harm from Language Models
- Bone Soups: A Seek-and-Soup Model Merging Approach for Controllable Multi-Objective Generation
- Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
- Preference Learning Unlocks LLMs' Psycho-Counseling Skills
- SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks