Reinforcement Learning from Human Feedback

arXiv:2504.12501 · cs.LG · Submitted 2025-04-16 · Read on arXiv

cs.LG

Submitted: 2025-04-16

Updated: 2026-09-11

Code: https://github.com/tatsu-lab/stanford_alpaca

Project page: https://danieltakeshi.github.io/2017/04/02/notes-on-the-generalized-advantageestimation-paper

Terminology

Sources

Related papers