A Probabilistic Approach for Model Alignment with Human Comparisons
cs.LG, stat.ML
Submitted: 2024-03-16
Updated: 2026-09-24
Terminology
Sources
- Robust Variational Autoencoder
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Improving Human Sequential Decision-Making with Reinforcement Learning
- Algorithmic Decision-Making Safeguarded by Human Knowledge
- Efficient Exploration for LLMs
- Improving alignment of dialogue agents via targeted human judgements
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- Dialogue Learning With Human-In-The-Loop
- Towards an Automatic Turing Test: Learning to Evaluate Dialogue Responses
- Robust Feature Learning for Multi-Index Models in High Dimensions
- Aligning Model Properties via Conformal Risk Control
- Foundation Model's Embedded Representations May Detect Distribution Shift
- Shattering the Agent-Environment Interface for Fine-Tuning Inclusive Language Models
- LOLA: LLM-Assisted Online Learning Algorithm for Content Experiments
- Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons
- Fine-Tuning Language Models from Human Preferences
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks