The Weakest Link Tells It All: Outcome-Supervised Process Reward Modeling via Learnable Credit Assignment

arXiv:2606.27739 · cs.LG · Submitted 2026-06-26 · Read on arXiv

cs.LG

Submitted: 2026-06-26

Updated: 2026-09-26

Code: https://github.com/openai/prm800k

Terminology

Sources

Related papers