ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation

arXiv:2608.10905 · cs.LG · Submitted 2026-08-11 · Read on arXiv

Ximo Zhu, Ruiqi Liu, Rong Wang, Ping Wu, Xiang Zheng, Wenzhuo Xu, Xubin Yao, Zhiyuan Yan, Bo Li, Jun Gao, Xiaolei Lv

Hello Group Inc. · Institute of Automation, Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · Peking University

cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable.

Terminology

Summary

On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher–student agreement to weight, filter, or truncate the sampled trajectory. These signals do not directly determine whether the teacher can continue a student prefix to a correct answer, and trajectory-level interventions can conflate one rollout’s unreliability with low expected training value of its prompt. We define prompt-level teacher continuation reliability R as the teacher’s probability of reaching a correct answer from a student prefix, averaged over prefixes and trajectories induced by the current student. Oracle experiments show that high-R prompts yield larger OPD gains and that descending-R training outperforms random and ascending orders on a fixed prompt pool. Because estimating R requires many teacher continuations, we use the maximum ROUGE-5 F1 between one independent student rollout and verifier-correct same-prompt teacher trajectories. Across ten equal-frequency bins of this actual score, mean R rises monotonically, showing that the proxy separates coarse reliability levels. ReOrder-OPD sorts prompts by the proxy, then draws independent on-policy training trajectories for vanilla OPD. It improves every matched aggregate comparison across Qwen3 and Gemma4 mathematics settings and Qwen3 code settings. Gains in all six FiRe-OPD and ExOPD settings show that prompt ordering complements within-trajectory supervision.

Improvements for AI systems

Improvements to AI Systems:

  1. Reliability-Aware Curriculum Learning for Distillation: Replace uniform or random prompt sampling during on-policy distillation with a reliability-sorted curriculum. The AI system first computes a proxy for teacher continuation reliability (max ROUGE-5 F1 between a student rollout and verifier-correct teacher trajectories for the same prompt). It then trains on prompts in descending order of this proxy, ensuring the student first learns from prompts where the teacher is most likely to guide it to correct answers, then gradually moves to harder, less reliable prompts. This yields faster convergence and higher final accuracy on math and code generation tasks.

  2. Adaptive Prompt Selection in Online Distillation: During training, dynamically re-rank the prompt pool based on the student's current reliability scores (recomputed every few steps). The system prioritizes prompts with high current reliability for vanilla OPD updates, while deferring low-reliability prompts until the student improves. This prevents wasted compute on trajectories where teacher supervision is likely to be noisy or unhelpful, improving sample efficiency.

  3. Hybrid Supervision with Reliability Gating: Combine trajectory-level filtering (e.g., teacher-student agreement) with prompt-level reliability ordering. The improved system first sorts prompts by the ROUGE-5 proxy, then applies existing within-trajectory supervision methods (like FiRe-OPD or ExOPD) only on the top-reliability half of the prompt pool. This avoids conflating a single bad rollout with a low-value prompt, leading to more stable gradients and better generalization.

  4. Verifier-Corrected Teacher Trajectory Retrieval for Few-Shot Prompting: Use the same ROUGE-5 proxy to build a small set of exemplar prompts for few-shot prompting or in-context learning. The system retrieves prompts where the teacher reliably reaches correct answers, then uses those teacher trajectories as few-shot demonstrations for the student at inference time. This improves the student's zero-shot and few-shot performance on math word problems and code generation.

  5. Early Stopping and Compute Allocation: Use the reliability proxy as a stopping criterion. If the mean reliability across a prompt bin drops below a threshold, the system halts training on that bin and reallocates compute to higher-reliability bins. This reduces total training time while maintaining or improving final performance.

What the Improved AI System Can Do:

  • Achieve higher accuracy on mathematical reasoning and code generation benchmarks (e.g., GSM8K, MATH, HumanEval) compared to standard on-policy distillation, with the same or less training compute.

  • Adaptively focus learning on prompts where teacher guidance is most trustworthy, avoiding wasted updates on noisy trajectories.

  • Provide more reliable few-shot demonstrations by selecting prompts with verifiable teacher success, improving inference-time performance.

  • Scale to larger prompt pools by automatically prioritizing high-value prompts, making distillation feasible with limited compute budgets.

Sources

Related papers