Demystifying the Unreasonable Effectiveness of Greedy Alignment Methods
cs.LG, cs.AI, cs.CC, cs.CL
Submitted: 2026-04-19
Updated: 2026-09-26
Terminology
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Human Alignment of Large Language Models through Online Preference Optimisation
- How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative Dynamics
- Active Preference Optimization for Sample Efficient RLHF
- Simple Regret Minimization for Contextual Bandits
- Asymptotic Instance-Optimal Algorithms for Interactive Decision Making
- KTO: Model Alignment as Prospect Theoretic Optimization
- Direct Language Model Alignment from Online AI Feedback
- ORPO: Monolithic Preference Optimization without Reference Model
- Reinforcement Learning from Human Feedback with Active Queries
- Coverage Improvement and Fast Convergence of On-policy Preference Learning
- Proportional Response: Contextual Bandits for Simple and Cumulative Regret Minimization
- Provably Efficient Regularized Online RLHF with Generalized Bilinear Preferences
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- Nash Learning from Human Feedback
- Towards Understanding the Influence of Reward Margin on Preference Model Performance
- Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
- Learning to summarize from human feedback
- Secrets of RLHF in Large Language Models Part II: Reward Modeling
- Self-Play Preference Optimization for Language Model Alignment
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks