Reinforcement Learning from Human Feedback
cs.LG
Submitted: 2025-04-16
Updated: 2026-09-11
Code: https://github.com/tatsu-lab/stanford_alpaca
Project page: https://danieltakeshi.github.io/2017/04/02/notes-on-the-generalized-advantageestimation-paper
Terminology
Sources
- WebGPT: Browser-assisted question-answering with human feedback
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- A Long Way to Go: Investigating Length Correlations in RLHF
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- The Art of Scaling Reinforcement Learning Compute for LLMs
- Olmo 3
- A General Language Assistant as a Laboratory for Alignment
- Constitutional AI: Harmlessness from AI Feedback
- Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2
- The Llama 3 Herd of Models
- Nemotron-4 340B Technical Report
- Scalable agent alignment via reward modeling: a research direction
- Fine-Tuning Language Models from Human Preferences
- Recursively Summarizing Books with Human Feedback
- Teaching language models to support answers with verified quotes
- Improving alignment of dialogue agents via targeted human judgements
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Llama 2: Open Foundation and Fine-Tuned Chat Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks