H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation".
Jane: The paper was written by Qixiang Yin, Huanjin Yao, Cai Yuchen, Jianghao Chen, Ziyi Wang et al. from Beijing University of Posts and Telecommunications (BUPT) and ByteDance and University of Science and Technology of China (USTC) and Beijing Key Laboratory of Network System and Network Culture and Key Laboratory of Interactive Technology and Experience System, Ministry of Culture and Tourism and Zhongguancun Academy, Beijing, China.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
The Core Idea: Tom: So, Jane, before we get into the specifics of how it works, can you explain the core problem H-OPD is solving in simple terms?
Jane: Certainly. Most existing AI models rely on a fixed assignment—if it's an image question, use the Vision-Language teacher; if it’s just text, use the Text teacher. But H-OPD recognizes that this is too rigid for multimodal tasks.
Lu: Right, because visual perception and abstract reasoning don't happen at the same time in a single complex query. They happen sequentially along the response trajectory, so one static teacher isn's sufficient throughout the entire generation process.
Meng: That’s where it gets interesting—the core idea of exposing the same input to both teachers, which is a huge change from traditional methods like ExOPD. But how do they manage two different modality streams for one input?
Lalam: It suggests that an AI doesn't have to commit to a single "expertise" but rather must be able to dynamically blend perspectives, which is a powerful idea for improving the quality of the content we generate.
The Mechanism: Tom: That dynamic blending is key, and I want people to really understand *how* it works in H-OPD. It’s not just a mix; it’s very specific.
Jane: It uses two main mechanisms: first, the vision-to-language description transfer, and then this clever idea of confidence-aware arbitration.
Lu: The V-to-L transfer is genius because it allows the text-only teacher to gain access to visual semantics through a textual proxy, effectively merging the two fields without requiring a complex shared visual latent space.
Meng: And when they talk about confidence-aware arbitration, that sounds like they are using entropy—the measure of uncertainty—to decide how much weight to give each teacher at the token level.
Lalam: That focus on predictive entropy is so important because it means the AI isn't just blindly averaging results; it’s actively deciding which teacher is currently more reliable for a specific step in the process, making its outputs far more trustworthy.
The Results and Impact: Tom: Looking at the results, H-OPD seems to be outperforming everything else, especially when comparing it to existing on-policy distillation methods like GRPO.
Jane: It consistently outperforms them across all eleven benchmarks listed in Table one which is a massive indicator of success.
Lu: And I see the potential for even greater scalability; if this technique scales with larger models, the gains could be exponential.
Meng: The token efficiency improvements shown in Figure five suggest that from an engineering standpoint, we can get better results using fewer computational resources than previous methods.
Lalam: This efficiency is vital because it means that the high-quality reasoning provided by H-OPD can be deployed more widely and even serve global communities with greater reliability.
Conclusion: Tom: So, to wrap up, H-OPD has successfully moved us from rigid, sample-level routing to this fluid, token-level arbitration.
Jane: It's a robust framework that finally lets the visual grounding and abstract reasoning capabilities collaborate effectively inside the student model.
Lu: I think this work opens up so many new research avenues for how AI can truly reason, not just mimic patterns.
Meng: Practically, it means we’ can start building real-world systems that handle complex multimodal inputs with a level of consistency we've never seen before.
Lalam: H-OPD is a powerful tool for improving the overall quality and reliability of AI in service of humanity. Thank you for listening to us discussing H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation.
Qixiang Yin, Huanjin Yao, Cai Yuchen, Jianghao Chen, Ziyi Wang, Min Yang*, Fei Su, Zhicheng Zhao
Beijing University of Posts and Telecommunications (BUPT) · ByteDance · University of Science and Technology of China (USTC) · Beijing Key Laboratory of Network System and Network Culture · Key Laboratory of Interactive Technology and Experience System, Ministry of Culture and Tourism · Zhongguancun Academy, Beijing, China
cs.CV, cs.LG
Submitted: 2026-08-24
Updated: 2026-08-25
Code: https://github.com/buptyqx/H-OPD
Importance score: 82/100
The gist: This paper introduces H-OPD, a confidence-aware heterogeneous multi-teacher on-policy distillation framework designed to enhance multimodal reasoning in Multimodal Large Language Models (MLLMs).
Key concepts
- Rigid Assignment
- Traditional AI models use a fixed assignment, meaning they rely on a static teacher (e.g., Vision or Text) based on the input type. H-OPD moves past this rigidity by allowing the AI to dynamically blend different perspectives during the generation process.
- Confidence-Aware Arbitration
- This mechanism uses predictive entropy, which measures uncertainty, to decide how much weight to give each AI teacher at every token level. This allows the system to actively choose the most reliable teacher for a specific step in the process.
- V-to-L Transfer
- The V-to-L transfer is a method allowing a text-only AI model to understand visual semantics. It uses a textual proxy derived from an image, effectively merging the two fields without requiring complex shared internal data spaces.
Terminology
Summary
This paper introduces H-OPD, a confidence-aware heterogeneous multi-teacher on-policy distillation framework designed to enhance multimodal reasoning in Multimodal Large Language Models (MLLMs). It addresses the limitations of existing on-policy distillation (OPD) methods that rely on static teacher routing,
which fails to account for the fact that different decoding steps require varying levels of visual perception versus abstract reasoning. By enabling collaborative supervision between vision-language and text-only teachers, H-OPD aims to provide more fine-grained, token-state-aware supervision
for complex reasoning trajectories.
The Motivation for Token-Level Arbitration
Existing post-training frameworks often assign a single teacher to an entire input instance based on task type, assuming that the entire generation trajectory of a given input should be supervised by the same teacher.
However, multimodal reasoning is a dynamic process where perceptual information may be misinterpreted or overwritten during later reasoning process.
This suggests that teacher reliability fluctuates depending on the specific token being generated.
The authors argue that vision-language (VL) and text-only teachers are not redundant but provide fine-grained complementary supervision.
Specifically:
% The VL teacher is more effective at perception-heavy steps.
% The text-only teacher can be more reliable at reasoning-intensive stages.
To exploit this, H-OPD moves away from task or sample-level routing toward a paradigm of token-level teacher arbitration along the shared student trajectory,
allowing multiple teachers to contribute to the same reasoning process with varying authority.
The H-OPD Framework Architecture
H-OPD employs two primary technical innovations to facilitate this heterogeneous collaboration. First, it introduces vision-to-language description transfer
to bridge the modality gap. Since text-only teachers lack direct visual access, the framework uses a VL teacher to extract task-relevant visual semantics
into a textual proxy (an image description). This allows the text teacher to focus exclusively on logical reasoning while still being informed by the visual context.
Second, the framework implements a confidence-aware arbitration mechanism.
At each decoding step, H-OPD calculates the predictive entropy of both teachers to measure their uncertainty. The arbitration weight is determined by these confidence scores:
-
The predictive entropy of each teacher is used to calculate a confidence score.
-
These scores are transformed into weights using an exponential function and a temperature parameter.
-
The final
arbitrated target distribution
is a weighted combination of the two teachers' predictions over a merged vocabulary of top-k candidates.
Experimental Results and Performance
The researchers conducted extensive evaluations over 11 widely-used reasoning benchmarks,
including MathVista, LogicVista, and CharXiv. The results demonstrate that H-OPD consistently outperforms standard OPD and GRPO across diverse tasks. Key findings include:
% Higher training efficiency:
% Superior test reasoning accuracy:
% Effective scaling:
The study shows that H-OPD benefits from both stronger teachers and larger students.
Furthermore, qualitative analysis through a response heatmap
confirms that the model successfully assigns higher weights to the VL teacher during visual grounding steps and to the text teacher during logical reasoning steps, achieving a more faithful coupling between perception and reasoning.
Limitations
The authors identify several constraints of their approach. The effectiveness of the text-only teacher is contingent upon the quality of the transferred image description; if descriptions are incomplete, noisy, or weakly aligned,
supervision becomes suboptimal. Additionally, while using predictive entropy as a proxy for reliability is effective, it does not guarantee correctness if teachers are overconfident or poorly calibrated.
Finally, the method introduces additional training overhead due to querying multiple teachers during rollout.
Improvements for AI systems
To implement the methodologies proposed in this paper, I would transition current multimodal reasoning architectures from static, task-based routing to a dynamic, token-level arbitration framework.
Here are the specific improvements and their corresponding capabilities:
I would implement a two-stage training pipeline consisting of an offline vision-to-language description transfer followed by online on-policy distillation. This replaces standard Supervised Fine-Tuning (SFT) with a system that uses heterogeneous teachers (a Vision-Language teacher and a Text-only teacher) to provide dense, tokenized supervision.
The improved AI system will be able to:
-
Perform Token-Level Expertise Switching: Instead of relying on a single model for an entire response, the system will dynamically shift its
authority
at every decoding step. It will automatically prioritize the Vision-Language (VL) teacher during stages requiring precise visual grounding (e.g., identifying coordinates or geometric shapes) and switch to a Text-only teacher during stages requiring heavy abstract reasoning or logical deduction. -
Mitigate Perceptual Drift in Long-Chain Reasoning: By using the text-only teacher to supervise reasoning steps through a
textual proxy
of the image, the system will prevent visual information from being misinterpreted oroverwritten
by incorrect reasoning steps during long autoregressive generation. This ensures that logical errors do not propagate from misperceived visual cues. -
Achieve Higher Training Efficiency and Convergence: By utilizing a confidence-aware arbitration mechanism (based on predictive entropy), the system will focus its learning signal on the most reliable teacher at any given moment, significantly reducing the number of training tokens required to reach high accuracy compared to standard Reinforcement Learning (RL) or Group Relative Policy Optimization (GRPO).
-
Enable Collaborative Multimodal Supervision: The system will bridge the modality gap by allowing a text-only model to act as a specialized logic expert for multimodal tasks, using high-fidelity textual descriptions of visual semantics provided by the VL teacher. This allows for much finer-grained supervision than current methods that simply route samples to either a vision or text model.
Sources
- GPT-4 Technical Report
- Qwen3-VL Technical Report
- Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
- On Predictability of Reinforcement Learning Dynamics for Large Language Models
- HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- Reinforcement Learning via Self-Distillation
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
- From System 1 to System 2: A Survey of Reasoning Large Language Models
- MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
- Gemini: A Family of Highly Capable Multimodal Models
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- Perception-Aware Policy Optimization for Multimodal Reasoning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- MiMo-V2-Flash Technical Report
- LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
- Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models