Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective".
Jane: The paper was written by Feng Zhang, Xinhong Ma, Ziqiang Dong, Xi Leng, Jianfei Zhao et al. from Beijing Institute of Technology and Qwen Business Unit of Alibaba and The Chinese University of Hong Kong, Shenzhen and Zhongguancun Academy.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: So, when we talk about "Revisiting Reinforcement Learning with Verifiable Rewards," we are looking at a paradigm where the reward comes from an external check, like a human grading it, rather than just letting the AI decide what’s good. But this paper is not just tweaking that; it’s fundamentally changing how we view the reward structure.
Jane: It seems like the authors found that while GRPO—the current standard method—is useful, it lacks a granular way to assign credit or blame for specific steps in a rollout, which is what makes "verifiable rewards" so powerful here.
Lu: The idea of moving beyond just score maximization to a contrastive framework opens up possibilities where the AI could learn not just *what* the correct answer is, but *why* it is correct relative to every other path I could have taken.
Meng: If we can focus on that "relative" aspect, we can design training loops that actually penalize specific types of errors rather than just accepting a general score reduction across all methods.
Lalam: When this capability reaches widespread use, it suggests a shift in how people interact with AI—it moves from merely seeking an answer to engaging in a dynamic process of improvement and learning alongside the model.
Summary: Tom: The paper starts by summarizing how current RLVR techniques have tried to improve GRPO by adjusting loss functions or adding sampling methods, but it points out that these are often just band-aids. They haven't addressed the core optimization mechanism at all.
Jane: The authors found two main issues with the existing Group Relative Policy Optimization model: first, that it uses surrogate scores instead of actual generation likelihoods, and second, a problem with how credit is assigned to individual rollouts within a group.
Lu: That score-insensitive credit assignment is fascinating; it means the old system treats every single positive rollout in a group as equally good or bad, regardless of how close they are to the distractors.
Meng: That's a huge operational hurdle; if we can’t tell the difference between a positive that was barely better than five negatives and we could, this achieving high performance on specific hard problems becomes much harder to guarantee.
Lalam: When AI can distinguish subtle differences in its own reasoning—seeing which paths were nearly correct and which ones were truly wrong—it's moving closer to the sophisticated form of human understanding.
Improvements & Methodology: Tom: To fix these issues, ConSPO, the method they propose, uses a massive technical leap by replacing those old surrogate scores with length-normalized sequence log probabilities. This is a huge shift toward alignment with actual generation likelihood.
Jane: Think of it this as making sure the training score matches how the model actually generates text during inference; no more mismatch between what we train on and what we use in practice.
Meng: The practical benefit here is that when you are trying to fix a deployment, you need accurate metrics, and ConSPO provides those accurate likelihood-based scores instead of approximations.
Lu: Then the method uses an InfoNCE-style objective—a contrastive function—to actively push the positive rollouts away from the negative distractors within the same group.
Lalam: This means our AI won't just be optimizing for a general 'good' answer; it will be actively striving to excel in its ability to distinguish its best path from all of those alternatives, making it a much more discerning reasoning tool.
Conclusion: Tom: So, we’ve seen that "Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective" introduces ConSPO to fix the limitations of existing methods by aligning scores and using contrastive learning. It's truly a paradigm shift for LLM training.
Jane: The results are impressive, showing consistent improvements across various benchmarks and model sizes, which is a huge relief for any applied research in this field.
Lu: I can only hope that this opens the door to even more creative ways of training AI, not just better mathematical reasoning but complex problem-solving across disciplines.
Meng: It's reassuring to see that the practical impact extends across different model sizes and datasets; it suggests a robust solution that scales well in deployment.
Lalam: This improvement means our cultural interaction with AI will become more accurate and sophisticated, reflecting the ability to truly distinguish good reasoning from poor reasoning.
Tom: It's clear this paper is setting a new standard for how we train LLMs for verifiable, high-quality reasoning tasks. We have a lot of excitement about ConSPO and its findings.
Beijing Institute of Technology · Qwen Business Unit of Alibaba · The Chinese University of Hong Kong, Shenzhen · Zhongguancun Academy
cs.LG, cs.AI
Submitted: 2026-05-13
Updated: 2026-09-18
Importance score: 1/100
The gist: The paper, titled "Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective," details an advanced framework for online Reinforcement Learning from Verifiable Rewards
Key concepts
- Verifiable Rewards
- This concept involves the reward coming from an external check, such as human grading, rather than the AI deciding what is good. The paper focuses on moving beyond simple score maximization to a framework that allows the AI to learn why its correct answer relative to alternatives are important.
- GRPO
- Group Relative Policy Optimization is a current standard RL method used in LLM training. The hosts note it has limitations, including using surrogate scores instead of actual generation likelihoods and failing to assign credit accurately for specific steps in a rollout.
- ConSPO
- This is the proposed method that fixes issues with existing RL techniques. It replaces surrogate scores with length-normalized sequence log probabilities and uses an InfoNCE-style contrastive function to actively push positive outcomes away from negative distractors.
Terminology
Summary
The paper, titled Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective,
details an advanced framework for online Reinforcement Learning from Verifiable Rewards (RLVR) that leverages contrastive principles.
The theoretical underpinnings of the approach involve complex mathematical derivations related to gradient calculation within a contrastive objective function, J bNCE(q). The text presents several identities crucial to understanding the optimization process:
- Identity for P ij: A relationship is established for the probability term P ij:
P ij = 1 over N sum i=1 N (1 - P i+) - 1 over N sum i=1 N X i = -+ over−
- Derivative Identities: The paper derives the partial derivatives of the objective function with respect to rollout states s-j and s-k.
d J bNCE(q) over d s-j = d J bNCE(q) over d s-k
- Key Gradient Relationship (Equation 12): A critical identity is derived by combining the gradient relationships with a specific probability calculation:
d J bNCE(q) over d s-j - d J bNCE(q) over d s-k = (product i=1 N P ij - 1) - (product i=1 N P ik - 1)
This identity simplifies to:
d J bNCE(q) over d s-j - d J bNCE(q) over d s-k = e-s-j - s-k over tau (product i=1 N P ij - 1) - e-s-j - s-k over tau (product i=1 N P ik - 1)
The core contribution of the work lies in adapting the contrastive principle to online RLVR. The authors note that their method differs significantly from existing preference optimization techniques:
-
Unlike methods that optimize
policy likelihoods over preference pairs without explicit reward modeling,
or SimPO, which useslength-normalized sequence logprobabilities as an implicit ranking score,
ConSPO is unique because ittransfers the contrastive principle to online RLVR, where positive-negative comparisons are constructed from verifiable rewards on policy-generated rollouts rather than from static pairwise preference data.
-
The proposed framework optimizes a
group-wise InfoNCE-style objective for online RLVR, using likelihood-aligned sequence scores and a curriculum-scheduled margin to enable contrast-sensitive credit assignment within rollout groups.
The models are evaluated on seven established benchmarks designed to test mathematical reasoning ability:
-
AIME24–26: Problems from the American Invitational Mathematics Examination, requiring
precise multi-step reasoning rather than multiple-choice selection.
-
HMMT25: Problems from the Harvard–MIT Mathematics Tournament, covering areas like
algebra and number theory, combinatorics, and geometry.
-
MATH500: A 500-problem evaluation subset of the MATH benchmark, spanning seven subjects including
algebra, geometry, counting and probability, number theory, and precalculus.
-
AMC: Problems from the American Mathematics Competitions.
-
OlympiadBench: An Olympiad-level benchmark containing challenging mathematics and physics problems.
The evaluation aims to reduce variance caused by small test sizes by reporting avg@32 on AIME, HMMT, and AMC
and pass@1 on the remaining benchmarks.
The training and evaluation utilize specific configurations summarized in Tables 7 and 8. The primary models are R1-Distill-Qwen and R1-Distill-Llama.
-
Training Configurations (Table 7): These include settings for
max prompt length(1024),max response length(8192),rollout temperature(1.0), a learning rate of 2 times 10-6, and specific values for the contrastive temperature (tau=4), maximum margin (M=10), and margin warmup ratio (alpha=30%). -
Evaluation Configurations (Table 8): These set the
max response lengthto 8192, a rollout temperature of 0.6, and a top-p value of 0.95.
A unified prompt template is used for both rollout generation and evaluation across all methods, which instructs the model to solve the problem step by step and place the final answer within,
facilitating answer extraction and rule-based verification.
Improvements for AI systems
The following improvements detail the implementation of ConSPO (Contrastive Sequence-level Policy Optimization) to replace or augment standard GRPO in any existing LLM reinforcement learning framework.
The fundamental upgrade involves replacing the traditional Group Relative Policy Optimization (GRPO) objective, which suffers from likelihood misalignment and score-insensitive credit assignment, with a novel contrastive objective.
Improvement: Replace the clipped token-level importance sampling ratio scores (clip(rho, 1-epsilon, etc.) used in GRPO with length-normalized sequence log-probabilities.
- Technical Specification: For every generated rollout o and query q, the score is defined as:
s theta(o, q) = sum t=1 o pi theta(o t q, o<t)
- System Capability: This ensures that the training signal directly corresponds to the actual sequence probability that governs autoregressive generation. The model is no longer optimizing an abstract ratio; it is optimizing its ability to generate high-probability sequences, leading to a direct alignment between training objective improvement and inference quality.
Improvement: Replace the fixed, group-level weighting p(q)(1 - p(q)) used in GRPO with a dynamic, InfoNCE-style contrastive objective that leverages the current internal separation of positive and negative rollouts.
- Technical Specification: The optimization targets a group-wise InfoNCE loss:
L ConSPO(q) = -1 over N+ times N- times tau sum i=1 N+ e s theta(o+ i q) / tau - e s theta(o j q) / tau
where e times is the exponential function, tau is the contrastive temperature, and N+ and N- are the counts of positive and negative rollouts in a group.
- System Capability: This transforms credit assignment from a static, group-level calculation (which treats all positives equally) to an adaptive, rollout-level comparison. The system now assigns stronger updates based on the current separation between positive examples and high-scoring negative distractors, allowing it to focus on
hard negatives
that were previously ignored by GRPO.
Improvement: Implement a dynamically scheduled margin (m t) within the contrastive objective to manage training dynamics.
- Technical Specification: The effective score is adjusted using a margin m t derived from the training progress lambda t:
s'theta(o, q) = s theta(o, q) plus or minus m t
The margin m t starts at zero (allowing early coarse ordering) and gradually increases to a target maximum margin M as training progresses.
- System Capability: This mechanism prevents premature convergence or over-aggressive updates. It ensures that the system first learns general positive/negative distinction, and only later applies high-pressure separation requirements, resulting in more stable and robust optimization.
By implementing ConSPO, the improved AI system will achieve:
-
Superior Reasoning Performance: Consistently outperform existing state-of-the-art RLVR baselines on complex mathematical reasoning benchmarks (e.g., AIME, HMMT) across diverse model architectures and scales.
-
High Generalization: Demonstrate effectiveness regardless of the underlying LLM backbone (e.g., Qwen, DeepSeek, Llama) or the specific training dataset used for RLVR fine-tuning.
3 Enhanced Learning Stability: Achieve more stable and efficient policy optimization compared to traditional gradient-based methods by adapting credit assignment based on real-time internal group dynamics rather than fixed statistical averages.
Sources
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
- Soft Adaptive Policy Optimization
- The Llama 3 Herd of Models
- GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
- Representation Learning with Contrastive Predictive Coding
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models
- Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
- Towards Flash Thinking via Decoupled Advantage Policy Optimization
- Group Sequence Policy Optimization
- Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
- Qwen3 Technical Report
- ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning
- PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks