On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers
cs.LG, cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: EMNLP 2026 Findings
Code: https://github.com/vigneshprabhakar1998/emnlp-2026-artifact-release
License: http://creativecommons.org/licenses/by/4.0/
The gist: Compact instruction-following rerankers are attractive for deployment, but conventional distillation pipelines typically train students by offline imitation of teacher outputs on a fixed set of
Terminology
Abstract
Compact instruction-following rerankers are attractive for deployment, but conventional distillation pipelines typically train students by offline imitation of teacher outputs on a fixed set of examples, constraining supervision to the teacher's observed ranking space. We revisit reranker distillation through the lens of reinforcement learning. We propose a two-stage framework combining off-policy teacher optimization with on-policy student distillation. In Stage 1, a 4B teacher reranker is strengthened with off-policy GRPO using LLM-judge feedback on 88K instruction-following examples. In Stage 2, a compact 1B student samples rankings from its own policy and receives soft teacher-derived rewards on those rankings, coupling student exploration with knowledge transfer. Our strongest gains appear under distribution shift. On MAIR-11, the original 11-subset, 869-query evaluation, the proposed student reaches 0.7670 nDCG@6, outperforming offline listwise KD by +4.6 points. Controlled comparisons against offline pairwise RankNet KD and on-policy GKD show that neither changing the offline distillation objective nor moving teacher-distribution matching on-policy reproduces the performance of reward-based on-policy distillation over student-sampled rankings. The advantage persists on MAIR-Full: across all 126 tasks and 9,356 queries, the proposed method obtains the highest task-macro point estimates among the evaluated distillation variants, reaching 0.6808 nDCG@6 and 0.7865 MRR@6. It also exceeds two released 7B RL-trained rerankers on the comparable MAIR-11 evaluation, while the same Stage 2 training procedure consistently improves three architecturally distinct alternative student backbones. On the 9,861-query validation benchmark, the resulting 1B reranker achieves 0.7624 nDCG@6 while providing a favorable quality-efficiency tradeoff relative to larger alternatives.
Sources
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- Distilling the Knowledge in a Neural Network
- INSTRUCTIR: A Benchmark for Instruction Following of Information Retrieval Models
- RankVicuna: Zero-Shot Listwise Document Reranking with Open-Source Large Language Models
- RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze!
- Qwen3 Technical Report
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Zero-Shot Listwise Document Reranking with a Large Language Model
- Passage Re-ranking with BERT
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks