From Constitutions to Control: Interpretable Rewards for Aligning Language Models

arXiv:2609.33086 · cs.LG · Submitted 2026-09-27 · Read on arXiv

cs.LG

Submitted: 2026-09-27

Updated: 2026-09-27

Code: https://github.com/jgaeb/rubric-rewards

Terminology

Sources

Related papers