Aletheia: What Makes RLVR For Code Verifiers Tick?

arXiv:2601.12186 · cs.SE, cs.AI · Submitted 2026-08-17 · Read on arXiv

Vatsal Venkatkrishna, Indraneil Paul, Iryna Gurevych

INSAIT, Sofia University "St. Kliment Ohridski" · Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, Technical University of Darmstadt · National Research Center for Applied Cybersecurity ATHENE

cs.SE, cs.AI

Submitted: 2026-08-17

Updated: 2026-08-18

Comments: 31 pages, 6 figures

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 58/100

The gist: The paper introduces Aletheia, a controlled, execution-grounded testbed designed to facilitate a contamination-free analysis of code verifier training recipes across disparate model sizes and

Terminology

Summary

The paper introduces Aletheia, a controlled, execution-grounded testbed designed to facilitate a contamination-free analysis of code verifier training recipes across disparate model sizes and covariate shifts across two common verifier application scenarios. The authors state: "We introduce Aletheia, a controlled, execution-grounded testbed to facilitate a contamination-free analysis of code verifier training recipes across disparate model sizes and covariate shifts across two common verifier application scenarios."

The motivation for this work is that "Multi-domain thinking verifiers trained via Reinforcement Learning with Verifiable Rewards (RLVR) are a cornerstone of modern post-training. However, their adoption in code generation has lagged behind that of execution feedback due to the prohibitive costs of the full RLVR pipeline. The authors note that the lack of RLVR adoption for code verifiers is unsurprising because the full RLVR recipe is expensive and demands intricate orchestration of rollout, behavior, and reference policies."

The paper ablates three primary choices along the performance–cost tradeoff in RLVR: intermediate thinking traces, learning from negative samples, and on-policy training. The authors state: In this work, we ablate three primary choices along the performance–cost tradeoff in RLVR: intermediate thinking traces, learning from negative samples, and on-policy training.

The testbed is created through a four-stage pipeline: "(1) generating solutions for competition-level programming questions from CodeContests+ from a pool of Weak and Strong open-source LLMs; (2) obtaining ground-truth pass rates (PRs) for the obtained codes through execution using SandboxFusion; (3) constructing lists of 2-5 candidates where exactly one is fully correct. These lists are either Easy (second best PR < 0.5) or Hard (second best PR ∈ [0.7, 0.9]); and (4) partitioning the resulting data into completely disjoint training and evaluation sets across three covariate shifts: strong generators, hard comparisons, and adversarial prompts."

The evaluation datasets include: Aletheia-Heldout (in-distribution evaluation with Easy comparisons by Weak models), Aletheia-Strong (Easy comparisons by Strong models, testing robustness to a shift in generator capability), Aletheia-Hard (Hard comparisons generated by Weak models, testing easy-to-hard generalization), and Aletheia-Adv (adversarial robustness evaluation with modifications to codes based on biases in LLM judges).

The paper evaluates verifiers using two complementary metrics: ListAcc (average top-1 selection accuracy) as a predictor of downstream Best-of-N (BoN) performance, and Kendall's τ-b (Kτ) as a measure of the verifier's ability to reconstruct the full ranking of candidates, which better predicts downstream RL performance. The authors explain: "Despite having high correlation with BoN, accuracy alone is insufficient to predict the utility of a verifier as an RL reward-model... Even a perfectly accurate verifier can induce a flat loss landscape if it cannot sufficiently differentiate the relative quality of the incorrect codes."

The analysis validates findings across three model sizes (1.5B, 7B, and 14B), uncovering scale-dependent training dynamics. The key findings are:

RQ1 (Thinking traces): The contribution of thinking traces to verifier performance increases monotonically with model scale. For BoN, "GRPO-Instruct and GRPO-Think-4k differ by ≤ 2.4 BoN points on average for 1.5-7B verifiers, indicating the style of intermediate trace makes little difference at smaller scale. At 14B, the same comparison goes to 8.4 points. Training reasoning budget follows a similar pattern: for 1.5B, the 4k→8k doubling adds 4.8 BoN points, but 8k→16k adds only 1.8, indicating diminishing returns. In contrast, 7-14B models keep climbing up to 16k, with gains of 8.3 and 11.6 BoN points respectively at the 8k→16k step. Thinking traces are crucial for Easy-to-Hard generalization and Reasoning-style traces enable models to utilize increased inference compute."

RQ2 (On-policy learning): Off-policy training collapses below random at 1.5B, but is competitive with semi-online methods at larger sizes. The online-offline gap between DPO–GRPO narrows with scale but never closes: from 23.27 BoN points at 1.5B to 6.65 at 14B. The authors note that BO-GRPO recovers neither GRPO-Think's BoN performance nor offers a meaningful cost saving. Additionally, Inference-time scaling narrows the offline–online BoN gap, most visibly at 1.5B, but cannot close it.

RQ3 (Negatives): GRPO-Think outperforms RAFT at every scale with a near-constant gap of ≈12.6 BoN points. For RL performance, the Kτ gap grows from 7.2 at 1.5B to 10.8 at 14B, indicating a growing importance of negatives for RL performance. The authors find that RAFT becomes increasingly unstable at larger sizes, with performance even degrading over training and Negatives provide no discriminative benefit on Aletheia-Hard, with RAFT performing comparably to GRPO-Think.

The paper also conducts a Pareto optimality analysis with respect to cost and performance. The authors find that DPO-Think-14B occupies a unique position across all eight panels: at 7.09 per step, it achieves 14B-scale performance at a 5.2× lower cost than GRPO-Think-14B, placing it on the Pareto frontier across all evaluations. They further note: The GRPO-Instruct verifiers anchor the low-budget end, while RAFT is dominated on all BoN settings... GRPO-Instruct-7B has a good cost – performance tradeoff, and is on the Pareto frontier in all panels. For easy-to-hard generalization, "GRPO-Think earns its high cost by extending the BoN and RL frontiers in-distribution and in evaluations with stronger generators and adversarial perturbations. However, Aletheia-Hard is a notable exception, where GRPO-Think is dominated in both metrics."

The authors conclude: "Our analysis reveals that these components are synergistic, but the degree of their impact is scale-dependent: for both BoN and RL, on-policy learning is the primary performance driver for small verifiers, while thinking becomes the most vital factor as model size increases. Negatives consistently boost performance on top-1 selection, are monotonic contributors to RL performance, and prevent reward curves from degrading at larger sizes. They further state: We find that scaling inference-time compute with self-consistency yields only a minor performance boost in most cases and cannot compensate for the absence of any component. Ultimately, DPO-Think-14B is an optimal choice for training verifiers for all scenarios. Although the full RLVR recipe is more performant, its cost is justified only when shifts in generator capability and adversarial perturbations are expected. At low budgets, GRPO-Instruct-7B is a strong baseline and is optimal across all evaluations, similar to DPO-Think-14B."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system:

  • Implement: A training pipeline that automatically selects the optimal RLVR recipe based on model size:

  • For models ≤1.5B: Prioritize on-policy learning (GRPO) over thinking traces; use shorter reasoning budgets (≤8k tokens) to avoid diminishing returns.

  • For models ≥7B: Prioritize long thinking traces (up to 16k tokens) and negative samples; on-policy learning becomes less critical.

  • Result: The system can train verifiers that achieve near-peak performance at up to 5.2× lower cost by using DPO-Think for larger models instead of full GRPO.

  • Implement: A cost-performance optimizer that:

  • Uses DPO-Think-14B as the default recipe for general-purpose verification (Pareto-optimal across all evaluation scenarios).

  • Switches to GRPO-Think-16k only when peak performance is required under generator shifts or adversarial perturbations.

  • Uses GRPO-Instruct-7B for low-budget deployments.

  • Result: The system can deploy verifiers that match full RLVR performance at 1/5th the training cost in most scenarios.

  • Thinking Traces:

  • Implement adaptive reasoning budget scaling: 1.5B models saturate at 8k tokens, while 7-14B models benefit from up to 16k.

  • Use reasoning-style traces (with backtracking and subgoal-setting) specifically for easy-to-hard generalization tasks.

  • Negative Samples:

  • Always include negative samples for BoN selection (consistent 12.6-point gain across all sizes).

  • For RL reward modeling, increase negative sample importance with model scale (gap grows from 7.2 to 10.8 Kτ points from 1.5B to 14B).

  • Use negative samples to prevent training instability at larger scales (RAFT reward curves degrade without them).

  • On-Policy Learning:

  • For small models (<7B), use batch-online GRPO as a cost-effective alternative (within 6 points of full GRPO).

  • For large models (≥7B), prefer offline DPO-Think over batch-online methods (4.4× cheaper with comparable performance).

  • Implement: A three-axis evaluation protocol mirroring Aletheia:

  • Generator shift: Test against stronger generators than seen in training.

  • Hard comparisons: Include near-correct distractors (pass rates 0.7-0.9).

  • Adversarial perturbations: Apply the six most effective biases (authority, bandwagon, self-declared correctness, etc.).

  • Result: The system can identify verifiers that maintain performance under distribution shifts, avoiding brittle models that fail in real-world deployment.

  • Implement: Dual-objective optimization:

  • For BoN selection: Optimize top-1 ListAcc; negative samples provide consistent gains.

  • For RL reward modeling: Optimize Kendall's τ-b (full ranking reconstruction); on-policy training provides disproportionate benefits (9.0 vs 5.5 points over DPO-Think at 14B).

  • Result: The system can train verifiers optimized for their specific downstream use case, avoiding suboptimal trade-offs.

  • Implement: Self-consistency scaling (SC@K) with awareness that:

  • It provides only modest gains (cannot compensate for missing training components).

  • It works best with thinking-trace-trained models (flat curves for CoT-only models).

  • It can partially bridge the offline-online gap at small scales (DPO-Think-1.5B nearly matches BO-GRPO at K=8).

  • Result: The system can allocate inference compute efficiently, avoiding wasted resources on models that don't benefit from scaling.

  1. Train code verifiers that achieve 80.54% BoN accuracy and 52.03 Kτ at 14B scale, with the ability to select the optimal training recipe based on available compute and model size.

  2. Deploy verifiers that remain robust to:

  • Stronger generator models (only 5.2-point drop on Aletheia-Strong).

  • Adversarial code modifications (maintains performance with 16k thinking budget).

  • Near-correct distractors (improved via reasoning traces).

  1. Reduce training costs by up to 5.2× without significant performance loss by using DPO-Think-14B instead of full GRPO when peak robustness isn't required.

  2. Provide reliable reward signals for RL training by reconstructing full rankings (Kτ) rather than just selecting the best candidate, preventing flat loss landscapes and stalled learning.

  3. Avoid common failure modes:

  • Small verifier degeneration (DPO-Think-1.5B has only 43% parse rate).

  • Training instability at large scales (mitigated by negative samples).

  • Overfitting to easy examples (RAFT overfits without max-5 filtering).

  1. Make informed deployment decisions by evaluating verifiers on all three covariate shifts (generator capability, comparison difficulty, adversarial perturbations) before integration into downstream pipelines.

Abstract

Multi-domain thinking verifiers trained via Reinforcement Learning with Verifiable Rewards (RLVR) are a cornerstone of modern post-training. However, their adoption in code generation has lagged behind that of execution feedback due to the prohibitive costs of the full RLVR pipeline. In this work, we ablate three primary choices along the performance-cost trade-off in RLVR: intermediate thinking traces, learning from negative samples, and on-policy training. We introduce Aletheia, a controlled, execution-grounded testbed to facilitate a contamination-free analysis of code verifier training recipes across disparate model sizes and covariate shifts across two common verifier application scenarios. Our analysis reveals that the optimal training recipe is scale-dependent: on-policy learning is the primary performance driver for small verifiers, whereas the thinking budget becomes the most vital factor at larger scales. While leveraging negative samples has a consistent impact on top-1 selection accuracy across sizes, their contribution to ranking reconstruction increases monotonically with scale and plays a key role in stabilizing training at large sizes. Our Pareto optimality analysis demonstrates that eliminating on-policy training at larger model scales yields a verifier that performs comparably to the full RLVR recipe. Furthermore, we find that eschewing thinking traces serves as a compute-efficient strategy at lower budgets, offering a strong trade-off between training cost and verifier accuracy. Ultimately, our work provides the empirical foundation necessary to efficiently deploy robust code verifiers, thereby enabling their wider adoption in post-training pipelines for large code generation models.

Sources

Related papers