OVD: On-policy Verbal Distillation

arXiv:2601.21968 · cs.CL · Submitted 2026-01-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "OVD: On-policy Verbal Distillation".

Tom: On-policy Verbal Distillation (OVD) introduces a memory-efficient framework that transfers reasoning capabilities from large teacher models to smaller student models using discrete verbal feedback.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Moving on, let's talk about the title and who came up with this work. The full title is "OVD: On-policy Verbal Distillation," and it involves a team of researchers including Jing Xiong, Hui Shen, Shansan Gong, Yuxin Cheng, Jianghan Shen, Chaofan Tao, Haochen Tan, Haoli Bai, Lifeng Shang, and Ngai Wong.

Jane: That's quite a group of authors working on this project; it shows how collaborative these kinds of deep learning efforts really are. The title itself clearly signals the main innovation: using verbal distillation in an on-policy setting.

Lu: The names suggest a strong foundation in both large model architecture and reinforcement learning, which is exactly what you need to design something like OVD, as it blends those two fields so tightly.

Meng: I've seen papers where the author list points directly to their expertise; seeing this mix of deep learning and RL researchers suggests they have a solid grasp on both the model transfer and the training dynamics required here.

Lalam: The collaboration itself is interesting because it hints at how these specialized problems are being tackled by different teams, which can inspire other groups in developing novel distillation techniques.

The paper's summary: Tom: So, what's the actual essence of the OVD paper? Basically, they are proposing a new way to move reasoning power from a large teacher model to a smaller student model by ditching the need for token-level probability matching and instead using discrete verbal scores from teachers to guide their learning process.

Jane: That’s right, Tom. Instead of looking at every possible next word the teacher could pick, OVD uses those zero to nine verbal scores to evaluate if the student's reasoning path is good or bad at each stage. It essentially turns knowledge distillation into a trajectory matching problem instead of a probability matching problem.

Lu: That trajectory focus is key because it bypasses the need for massive logit storage, which they show requires about sixteen times more memory than the KV cache just for a single teacher model with thirty-two rollouts.

Meng: That memory reduction is what makes it practical; if you can reduce the complexity from something that scales linearly with sequence length and vocabulary size down to a structure dependent on reasoning steps and verbal vocabulary size, then we're talking about real deployment possibilities.

Lalam: The paper emphasizes that this method lets the student model freely explore the output space because it isn't strictly constrained by matching every single logit value, which is a major win for exploration.

The paper's improvements: Tom: Now let’s look at the specific improvements they claim OVD offers over what’s currently out there. They highlight several key things, most notably the massive memory reduction, but also significant performance gains on both Web Q andA and mathematical reasoning benchmarks.

Jane: The memory efficiency is huge; they state that by replacing token-level logits with verbal scores, the memory cost is reduced by a factor of approximately N times V over the old method. That factor really shows how much overhead was being cut down.

Lu: And they also achieved performance gains that are quite substantial; for instance, on math benchmarks, they showed a gain of up to twenty-five point seven percent when training with only one random sample per problem. That level of improvement on math tasks is pretty impressive considering the constraints they put on the training data.

Meng: The paper also points out that this method allows for better credit assignment for multi-step reasoning because it uses step-level verbal supervision instead of just an answer correctness reward, which is a practical win for debugging complex AI behavior.

Lalam: I also see the improvement in how the model learns to navigate its environment through verbal feedback; it’s not just learning the final result, but learning how to get there coherently across multiple steps because of that step-level supervision.

Conclusion: Tom: Alright, we’ve covered a lot regarding "OVD: On-policy Verbal Distillation," and to wrap up, the authors are really emphasizing how this framework successfully balances memory efficiency with high performance across different reasoning tasks. The main message is that trajectory matching guided by verbal feedback is a much smarter way to transfer complex AI capabilities.

Jane: It seems like the paper proves that you don't need an exact match on every token to achieve excellent student performance, especially when you use discrete scores from a teacher model as guidance. This makes the entire process more accessible for smaller models.

Lu: I think the real implication is that this opens up possibilities for creating highly capable reasoning agents that are deployed on edge devices or in environments with limited computational resources because of how efficiently they manage their memory footprint.

Meng: I see the immediate practical impact being in the training pipeline itself; if we can train these models more sample-efficiently, it means less time and fewer expensive runs needed to get a capable reasoning agent into production.

Lalam: From a cultural perspective, I think this points toward AI systems that are not just smart in isolation but are also coherent and methodical in their thinking, which is really valuable for building trust with users.

Jing Xiong, Hui Shen, Shansan Gong, Yuxin Cheng, Jianghan Shen, Chaofan Tao, Haochen Tan

cs.CL

Submitted: 2026-01-29

Updated: 2026-09-28

Project page: https://ovd.github.io

Importance score: 90/100

The gist: On-policy Verbal Distillation (OVD) introduces a memory-efficient framework that transfers reasoning capabilities from large teacher models to smaller student models using discrete verbal feedback.

Key concepts

Trajectory Matching
Instead of comparing every single word (token) the teacher predicts, OVD compares entire sequences of reasoning steps. This focuses the distillation on the overall coherence and correctness of a solution path rather than exact word-for-word matching, making it much more memory efficient.
Verbal Rejection Sampling
This technique selects high-quality training examples by sampling discrete scores (0–9) from the teacher's distribution. Trajectories with low scores are discarded, and the process is repeated until a good sample is found, ensuring only valuable reasoning steps are used for training.
On-policy Sampling
The student model generates its own data (trajectories) based on its current policy. This ensures the feedback it receives accurately reflects how *it* would perform, which helps prevent distribution mismatch errors common in other distillation methods.

Terminology

Summary

On-policy Verbal Distillation (OVD) introduces a memory-efficient framework that transfers reasoning capabilities from large teacher models to smaller student models using discrete verbal feedback. This method addresses the severe memory bottlenecks and exploration constraints found in traditional token-level distillation by replacing probability matching with trajectory matching guided by teacher verbal scores.

How it works

OVD reformulates knowledge distillation as trajectory matching rather than token-level probability matching, which dramatically reduces memory consumption. Instead of requiring the teacher model to output full vocabulary logits at every decoding step, OVD utilizes discrete verbal scores (0–9) from teacher models to evaluate reasoning correctness and coherence across reasoning steps. This approach replaces the prohibitive memory cost of storing logits over the entire vocabulary with a more efficient structure, reducing memory overhead from a term scaling linearly with sequence length (O(N · L · V)) to O(N · K · v), where K is the number of reasoning steps and v is the verbal vocabulary size.

The framework operates through several key components:

  1. Teacher Models: Tasks are categorized into Web Q&A (using an Environment Agent) and mathematical reasoning (using a Reasoning Agent). The Environment Agent encapsulates web Q&A tasks with verifiable outcomes, while the Reasoning Agent, such as a large model like QwQ-32B, provides fine-grained verbal feedback on reasoning correctness, coherence, and progress.

  2. On-policy Sampling: The student model generates trajectories from its own policy to ensure feedback on the student’s own distribution, mitigating distribution mismatch inherent in off-policy methods.

  3. Verbal Rejection Sampling: This mechanism is used to select high-quality trajectories for training. Instead of relying on inaccessible full-vocabulary logits, OVD operates over a compact score space by sampling discrete scores from the teacher's distribution and rejecting trajectories with low scores (Trajectories receiving low scores are rejected; for each rejected trajectory, we resample from the teacher agent and continue generation).

Key Contributions and Theoretical Guarantees

The paper makes several significant contributions regarding memory efficiency, performance, and theoretical rigor:

(i) Memory Efficiency:

OVD dramatically reduces memory consumption by replacing token-level logits with verbal scores. This enables longer trajectories, larger batch sizes, and more samples per problem without requiring teacher token-level distributions. The memory cost is reduced by a factor of approximately N · V /v.

(ii) Performance and Exploration:

OVD demonstrates substantial performance gains on Web Q&A and mathematical reasoning benchmarks. Experiments show "OVD substantially outperforms existing methods, delivering up to +12.9% absolute improvement in average EM on Web Q&A tasks and a up to +25.7% gain on math benchmarks (when trained with only one random samples). The method avoids token-level alignment, allowing the student model to freely explore the output space."

(iii) Theoretical Analysis:

The framework is analyzed through the lens of interactive imitation learning and verbal rejection sampling. The paper provides a formal proof showing that the procedure yields an unbiased gradient estimator (Theorem 3.1) under a mixture training distribution, and demonstrates variance reduction (Proposition 3.2) by replacing rejected trajectories with teacher demonstrations, leading to a bound where V[ˆgRS] ≤ V[ˆg0] − E[1[S(y) < θ]∥R(y)∇θ log πS(y)] + O(VT).

Training Procedure and Reward Design

The complete training pipeline integrates verbal distillation with Group Relative Optimization (GRPO). The reward design is outcome-based:

  1. Web Q&A uses the F1 score: R(y) = F1(y, y∗), measuring word overlap.

  2. Math reasoning uses a binary indicator: δ(y, y∗) ∈ [0, 1], denoting exact matching under mathematical equivalence (numerical tolerance).

The policy gradient objective is decomposed to allow for flexible credit assignment:

L = −Ex,y "X K k=1 Rk · log πS(skx, s<k)"

This formulation enables step-level scoring or trajectory-level scoring by setting Rk appropriately. Policy optimization uses GRPO with clipped gradients (PPO) and group relative advantages to normalize rewards within each problem's trajectory group.

Experimental Validation

Extensive experiments on Web Q&A and mathematical reasoning benchmarks confirm OVD’s superiority:

(i) Web Q&A:

OVD consistently outperforms search-based baselines like Search-o1, Search-R1, and ZeroSearch. For instance, on Qwen-2.5-3B-Base, OVD exceeds Search-R1 by 10.8 points (43.6% vs.

Improvements for AI systems

As a fastidious researcher, I have thoroughly analyzed the On-policy Verbal Distillation (OVD) paper. The core innovation is replacing memory-intensive token-level probability matching with memory-efficient trajectory matching using discrete verbal scores (0–9) from teacher models, integrated within an on-policy reinforcement learning framework.

Based on this scientific foundation, here are the specific improvements and capabilities for AI systems:


)

The improved AI system will possess the following capabilities:

Ability to perform long-horizon, complex reasoning tasks (e.g., multi-step mathematical proofs or intricate web Q&A chains) with significantly reduced memory overhead compared to traditional token-level knowledge distillation methods. This is achieved by scaling memory complexity from linear dependence on sequence length and vocabulary size to a much more manageable dependency on the number of reasoning steps and the verbal vocabulary size, enabling training on much longer trajectories.

Enhanced ability to leverage interactive environment feedback for reasoning refinement. The system can utilize Environment Agents (simulated search engines or specialized tools) as teachers to provide step-level quality assessments (verbal scores 1–10) on intermediate reasoning steps, allowing the student model to learn not just correct answers, but also the process of high-quality reasoning and tool use.

Robustness against distribution shift during distillation and improved exploration capabilities. By employing Verbal Rejection Sampling guided by these discrete scores, the system can efficiently filter out low-quality exploratory trajectories while retaining high-reward student samples. This mechanism ensures that the student model is trained on a mixture of its own successful explorations and expert teacher demonstrations, preventing catastrophic forgetting or mode collapse that plagues pure imitation learning methods.

Superior sample efficiency in RL training scenarios. The system demonstrates the ability to achieve substantial performance gains (up to +25.7% on math benchmarks) with very few random samples (e.g., one random example per problem). This makes it highly effective for deploying reasoning capabilities into models that are resource-constrained or have limited access to large, pre-labeled datasets.

Adaptive and self-regulating learning curriculum. The system inherently learns an adaptive curriculum: initially relying heavily on teacher demonstrations (low acceptance rate), and gradually shifting towards autonomous exploration as the student's policy improves (acceptance rate increases). This allows the model to transition smoothly from guided imitation to independent, high-quality reasoning behavior.

Precise quality assessment via granular feedback. The use of a discrete verbal vocabulary of size 10 allows for nuanced, qualitative judgments (e.g., distinguishing between partially correct and correct and well-reasoned). This fine granularity leads to a more precise trajectory selection mechanism, ensuring that the distillation process focuses on learning the nuances of quality rather than just binary correctness.

Versatility across reasoning domains. The framework is designed to be task-agnostic, adapting its trajectory generation structure for both sequential mathematical derivations and multi-hop web Q&A (which includes query generation and document retrieval steps), making it applicable to a wide variety of complex language tasks.

Sources

Related papers