Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
cs.LG, cs.AI, cs.CL, stat.ML
Submitted: 2026-04-06
Updated: 2026-09-26
Code: https://github.com/takahashihiroshi/rdro
Terminology
Sources
- GPT-4 Technical Report
- Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
- The Llama 3 Herd of Models
- Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model
- Binary Classifier Optimization for Large Language Model Alignment
- Preference Optimization by Estimating the Ratio of the Data Distribution
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Decoupled Weight Decay Regularization
- Nash Learning from Human Feedback
- The Woman Worked as a Babysitter: On Biases in Language Generation
- Qwen2 Technical Report
- Solving math word problems with process- and outcome-based feedback
- Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
- Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks