VERPO: Verified Evidence Regularized Policy Optimization
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-05
Updated: 2026-09-23
Comments: 38 pages, 9 figures, including appendices
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Verifiable outcome rewards guide language-model post-training, but sequence-level advantages do not identify which token-level decisions should be preserved or revised.
Terminology
Abstract
Verifiable outcome rewards guide language-model post-training, but sequence-level advantages do not identify which token-level decisions should be preserved or revised. Evidence-conditioned Teachers provide denser supervision by replaying sampled trajectories with privileged feedback. Yet indiscriminate imitation risks transferring formatting or reasoning-style shifts that do not support task success. We introduce VERPO, a Verified Evidence Regularized Policy Optimization framework that treats evidence as a proposal for policy correction while retaining the outcome objective. It separates evidence-free reference restoration from signed token-level evidence corrections. Fisher Evidence Contrast attenuates corrections along an estimated evidence-presence direction. A stopped token-wise ZPD controller scales acceptance according to local reward alignment and Fisher movement cost, while the reference channel remains independent of acceptance. Across five scientific-reasoning and tool-use tasks, the best variant on each backbone exceeds the strongest compared baseline in average score. The averages rise from 0.6826 to 0.6857 on Qwen3-4B, from 0.6895 to 0.7058 on Qwen3-8B, and from 0.4751 to 0.5657 on Llama-3.2-1B.
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models
- MiniLLM: On-Policy Distillation of Large Language Models
- Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation
- CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
- Distilling the Knowledge in a Neural Network
- Reinforcement Learning via Self-Distillation
- When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
- Rethinking On-Policy Self-Distillation for Thinking Models
- EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation
- Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
- Decoupled Weight Decay Regularization
- RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
- EviSD: Evidence-Conditioned Self-Distillation for Search-Augmented Agents
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks