Advantage Scale Calibration Imbalance in Group-Relative Optimization under Low-Variance Rewards: Diagnosis and Bounded Recovery

arXiv:2609.19164 · cs.CL, cs.LG · Submitted 2026-08-27 · Read on arXiv

cs.CL, cs.LG

Submitted: 2026-08-27

Updated: 2026-08-27

Comments: 10 pages,2 figures

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: In verifier-style RLVR, group-relative optimization often treats advantage scale as an implementation detail.

Terminology

Abstract

In verifier-style RLVR, group-relative optimization often treats advantage scale as an implementation detail. This paper separates two low-variance cases: sub-resolution jitter that should not become a preference signal, and credible but small cardinal gaps that should be learned without distorting KL calibration. We propose an advantage-scale three-way calibration interface: the same within-group scale denominator simultaneously determines the reward-branch strength, prompt-level batch weight, and the effective KL calibration induced when the reward branch is re-expressed on the original cardinal scale. This interface explains why RLOO / Dr.GRPO can let credible small gaps become KL dominated, whereas GRPO's standard-deviation denominator can amplify tiny gaps without bound. Based on this interface, we further introduce the Reward-Resolution Protocol and MaxNorm-AC, respectively filtering sub-resolution gaps and providing bounded cardinal recovery on credible nonzero gaps. Across dense / MoE architectures and math / code reasoning, MaxNorm-AC improves over the strongest robust-scale baseline while truncating the low-variance inverse-scale tail.

Related papers