Improving Reasoning Ability via Asynchronous On-Policy Self-Distillation under Positive Rollouts

arXiv:2605.06650 · cs.CL · Submitted 2026-05-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Improving Reasoning Ability via Asynchronous On-Policy Self-Distillation under Positive Rollouts".

Jane: Reinforcement learning with verifiable rewards (RLVR) is being advanced by Positive-Only Policy Optimization (POPO), a novel framework that enables policy improvement through online positive rollouts without relying on negative rollouts.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Moving on from what we just touched upon, let's look at the specific title and authors of this paper, "Improving Reasoning Ability via Asynchronous On-Policy Self-Distillation under Positive Rollouts." The title itself really highlights the core mechanism they are proposing: using positive rollouts to improve reasoning ability through self-distillation.

Jane: That title makes sense when you think about how the learning process is structured; it suggests a feedback loop where the model learns by comparing its output against successful examples without needing explicit negative feedback.

Lu: The authors, Mingwei Xu and Hao Fang from the University of Washington, are clearly deep in this RLVR space, and their work on asynchronous on-policy self-distillation shows a sophisticated way to manage that learning process across different policy updates.

Meng: It’s interesting how they framed it as asynchronous; that implies the system can handle multiple policy iterations or rollouts in parallel, which points toward potentially faster convergence times during training.

Lalam: I think the authors are really emphasizing the 'Positive-Only' part because it sets them apart from methods that still rely heavily on negative sampling, which is a key limitation they are trying to overcome.

The paper's summary: Tom: So, in terms of what the paper actually summarizes, it’s proposing a framework called Positive-Only Policy Optimization or POPO that learns exclusively through online positive rollouts. This system aims to enhance reasoning ability by focusing entirely on reinforcing correct responses during the learning phase.

Jane: It sounds like they are suggesting that since verifying rewards is deterministic, we can bypass the usual difficulty of getting meaningful negative signals and instead use a carefully constructed positive reinforcement signal to steer the policy toward better reasoning chains.

Lu: The summary points out that POPO uses bounded importance sampling over the positive rollout set to preferentially reinforce those successful responses, which is a specific mathematical technique they developed for this purpose.

Meng: I’m looking at that part about using bounded importance sampling; from an engineering view, normalizing over the positive set helps keep the reinforcement signal stable and prevents runaway gradients when we only have positive data available.

Lalam: It makes sense that they summarize it this way because it clearly contrasts their method against methods like GRPO, which relies on both positive and negative rollouts, showing why their approach is different.

The paper's improvements: Tom: Now let's talk about the specific improvements the paper details for this "Improving Reasoning Ability via Asynchronous On-Policy Self-Distillation under Positive Rollouts" framework. They introduce several components designed to make this positive-only approach work robustly.

Jane: The main improvement seems to be their structure, where they define a loss function that explicitly focuses on the positive set, and they use bounded importance sampling weights to create a self-competition situation among the correct answers.

Lu: I think their introduction of the siamese policy network with an EMA update law for a stabilized policy anchor is really clever because it provides that necessary stability against catastrophic drift that often happens when you only feed positive data.

Meng: That EMA anchoring sounds like a practical safeguard; if we can maintain a stable reference point for good reasoning patterns while the main policy evolves, it should make training much more predictable on real-world datasets.

Lalam: The representation-space alignment through a bounded similarity penalty instead of just KL divergence is another improvement that ensures the optimized policy stays semantically structured, which is crucial for coherent long-form outputs.

Conclusion: Tom: So, to wrap up our discussion on this paper, the authors conclude by summarizing how their POPO framework achieves performance comparable to or even better than GRPO in specific benchmarks like AIME two thousand twenty-five showing that positive-only optimization can work effectively for reasoning <ref:2605.06650#pg2>.

Jane: They emphasize that this methodology has implications for future sparse RLVR beyond relying on negative rollouts, suggesting a more scalable direction for enhancing LLM reasoning.

Lu: I think the overall implication is that we might be able to achieve higher levels of reasoning ability in models without needing the immense computational cost or sampling difficulty associated with generating high-quality negative examples.

Meng: For practical application, this means we can deploy these models faster because the training phase doesn't require us to spend significant time collecting and processing negative verification data.

Lalam: I really see the biggest impact here being in how we build AI culture; if we can optimize models solely on positive reinforcement, it shifts our focus from just correcting errors to actively cultivating superior reasoning skills within the model itself.

University of Washington

cs.CL

Submitted: 2026-05-07

Updated: 2026-10-02

Importance score: 93/100

The gist: Reinforcement learning with verifiable rewards (RLVR) is being advanced by Positive-Only Policy Optimization (POPO), a novel framework that enables policy improvement through online positive rollouts

Key concepts

Positive-Only Policy Optimization (POPO)
A novel RL framework that learns policy improvement exclusively from correct responses (positive rollouts). It avoids using negative rollouts, instead relying on mechanisms like probability redistribution and entropy loss to implicitly penalize wrong answers while maintaining stability.
Bounded Importance Sampling Weights
These weights normalize the probability of positive responses over the set of all positive outcomes. By using these weights, POPO preferentially reinforces confident correct answers while ensuring diversity through normalization, creating a self-competition effect.
Representation-Space Alignment
Instead of traditional KL divergence, POPO uses a similarity penalty in the representation space. This forces the policy network's embeddings to align with those of a stabilized Siamese network, ensuring semantic consistency without relying on explicit divergence measures.

Terminology

Summary

Reinforcement learning with verifiable rewards (RLVR) is being advanced by Positive-Only Policy Optimization (POPO), a novel framework that enables policy improvement through online positive rollouts without relying on negative rollouts. This approach addresses the limitation of sparse binary rewards by demonstrating that implicit negative gradients can emerge naturally through reinforcing positive probability via rollout redistribution, leading to performance comparable to or superior to existing methods like GRPO across various mathematical reasoning benchmarks.

The gist

We propose Positive-Only Policy Optimization (POPO), a novel RLVR framework in which learning can occur exclusively via online positive rollouts.

Core Mechanism of POPO

The POPO framework is designed to learn exclusively from correct responses while maintaining policy stability through several innovative mechanisms. The primary optimization objective, defined by the loss function LPOPO(θ), is structured to reinforce positive samples while incorporating regularization terms. Specifically, the loss function is given by:

LPOPO(θ) = −Ex∼D [∑ y∈S+(x) wθ (y x) · log πθ (y x)] + αLsim(θ, ϕ, ξ) + βLent(θ).

This loss function explicitly focuses the expectation only over responses in the positive set S+(x), which is defined as those responses where R(x, y) = 1.

Key Components for Policy Optimization

POPO introduces three distinct components to achieve its goals:

  1. Bounded Importance Sampling Weights: POPO utilizes bounded importance sampling weights, which normalize over the positive rollout set, to preferentially reinforce positive responses. These weights are defined as wθ (y x) = πθ (y x) / Z+(x), where Z+(x) is the sum of probabilities over the positive set. This scheme creates a self-competition situation, reinforcing confident correct answers while maintaining diversity through normalization.

  2. Adaptive Anchor: To stabilize policy evolution, POPO employs a siamese policy network with an exponential moving average (EMA) update law for a stabilized policy anchor. The parameters of this Siamese network (ξ) are updated using the formula: ξ ← τ · ξ + (1 − τ) · θ, where τ is the momentum coefficient.

  3. Representation-Space Alignment: Instead of relying on KL divergence, POPO replaces it with a bounded similarity penalty term in the siamese representation space. The loss term Lsim(θ, ϕ, ξ) is defined as: Lsim(θ, ϕ, ξ) = −Ex∼D [∑ y∈S+(x) wθ (y x) · coshϕ(fθ (x, y)), sg(fξ (x, y) + ϵ)]. This ensures the policy network should align with the Siamese network within the semantic embedding space.

Implicit Negative Gradients

A central finding of POPO is that it achieves stability and implicitly penalizes incorrect responses without using explicit negative rollouts. The paper proves this through two reinforcing forces:

  1. Weight Probability Redistribution: The softmax normalization over the positive set implicitly forces πθ (y x) to decrease for y′ ∈ S−(x). This is formalized in Theorem 3.1, showing that the gradient of LPOPO with respect to incorrect logits zy′ is strictly positive when it exceeds a specific threshold.

  2. Entropy Regularization: The entropy loss term, Lent(θ) = −H(πθ (· x)), adds a penalty that is strongest on the most probable incorrect responses.

Stability and Robustness

The framework's stability is further guaranteed by the asymmetric architecture and momentum adaptation. Lemma 3.2 establishes a bound on the parameter divergence between the policy network θt and the Siamese network ξt, showing that in steady state:∥θt − ξt∥ ≤ τ η Gmax / (1 − τ). This guarantees a bounded gap even without conventional KL divergence, preventing pathological gradient explosions. The similarity loss Lsim also provides a second restoring force, complementing the EMA anchoring.

Experimental Validation

Extensive experiments using publicly available, well-established text-LLM models, such as the Qwen family and DeepSeek-R1 distilled models, were conducted across all-level mathematical benchmarks including MATH-500, AMC23, AIME 2024/2025, and Olympiad. The results demonstrate that POPO achieves performance comparable to, or even superior to GRPO, notably achieving 36.67% in AIME 2025 with Qwen-Math-7B, outperforming GRPO's 30.00%. Ablation studies confirmed the necessity of components like importance sampling and momentum adaptation for robust performance.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the provided paper, Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients (POPO). The core innovation is developing a novel Reinforcement Learning from Verifiable Rewards (RLVR) framework that optimizes Large Language Models (LLMs) using only online positive rollouts.

Here are the specific improvements and capabilities this POPO framework enables in AI systems:


  1. 】Elimination of Negative Rollout Dependency for Policy Optimization

This is the most significant architectural improvement.

  • The system can perform policy post-training without relying on sampling or explicitly penalizing incorrect responses (negative rollouts). This bypasses the combinatorial vastness problem where penalizing a few negatives is unlikely to cover all failure modes.

  • The improved AI system can achieve stable policy optimization for LLMs, particularly in complex reasoning tasks, by focusing exclusively on reinforcing successful reasoning chains.

  1. 】Implicit Negative Gradient Generation via Weight Redistribution

Instead of explicit penalty terms (like negative rewards or KL divergence penalties), the system leverages the softmax normalization inherent in positive reinforcement to generate an implicit negative gradient signal.

  • The AI system can learn to tax incorrect responses by forcing their probability mass down relative to correct ones, even without a dedicated loss term for negatives. This allows for self-correction during policy updates.

3.】Stabilized Policy Evolution through Siamese Network Anchoring

The introduction of a siamese policy network with an Exponential Moving Average (EMA) adaptation law provides enhanced stability during the optimization process.

  • The improved AI system can maintain a stable anchor representation of good reasoning patterns while the main policy network learns, preventing catastrophic policy drift and gradient explosions.

4.】Representation-Space Alignment via Similarity Penalty

Replacing traditional KL divergence with a bounded similarity penalty in the Siamese representation space ensures that the model's internal logic remains semantically structured.

  • The AI system can ensure that the optimized policy maintains strong alignment with established reasoning structures in a latent space, leading to more robust and coherent outputs, rather than just maximizing raw log-likelihood.

5.】Improved Reasoning Performance on Sparse Binary Reward Tasks

The framework demonstrates superior performance compared to existing methods (like GRPO) on sparse binary reward benchmarks (e.g., AIME 2025).

  • The improved AI system can achieve state-of-the-art reasoning scores (e.g., reaching 36.67% in AIME 2025 with Qwen-Math-7B) on high-level mathematical and competitive exams, indicating a significant boost in complex problem solving capability.

6.】Model Agnostic and Scalable Application

The POPO framework is shown to be effective across various model architectures (Qwen family, DeepSeek models, Llama 3.1).

  • The improved AI system can be effectively deployed on diverse LLM backbones without extensive task-specific fine-tuning of the RL component itself, making it highly versatile for different domains (e.g., code generation, multimodal reasoning).

7.】Enhanced Robustness Against Overfitting and Collapse

The combination of bounded importance sampling weights and the asymmetric architecture prevents the policy network from collapsing into a degenerate solution (a risk in symmetric learning methods).

  • The AI system is more resilient to overfitting on the limited positive training data, ensuring that learned reasoning skills are generalizable rather than model-specific artifacts.

In summary, this paper proposes an RLVR paradigm shift: moving from positive and negative reinforcement to a positive-only self-reinforcing loop. The resulting AI system will be characterized by:

  • Higher accuracy in complex reasoning benchmarks.

  • Greater training stability and reduced hyperparameter sensitivity.

  • A more efficient use of computational resources during policy post-training, as it avoids the need for computationally expensive negative rollout sampling/filtering.

Sources

Related papers