Preference Data Selection for Mitigating the Alignment Tax in Large Language Models
cs.AI, cs.CL
Submitted: 2026-08-25
Updated: 2026-08-25
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- Evaluating Large Language Models Trained on Code
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Qwen2.5-Coder Technical Report
- AI Alignment: A Comprehensive Survey
- Mistral 7B
- Challenges and Applications of Large Language Models
- Overcoming catastrophic forgetting in neural networks
- Mitigating the Alignment Tax of RLHF
- Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
- Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
- The Llama 3 Herd of Models
- Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
- An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Large Language Model Alignment: A Survey
- Self-Distillation Enables Continual Learning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection