PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization

arXiv:2603.26535 · cs.AI · Submitted 2026-03-27 · Read on arXiv

cs.AI

Submitted: 2026-03-27

Updated: 2026-08-28

Code: https://github.com/project-numina/aimo-progress-prize

Terminology

Sources

Related papers