OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
National University of Singapore · Ludwig Maximilian University of Munich · Munich Center for Machine Learning · Sun Yat-sen University
cs.CV, cs.AI
Submitted: 2026-08-05
Comments: Corrected the uploaded manuscript. Project Page:https://github.com/aniri15/OPD-V
Code: https://github.com/aniri15/OPD-V
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 80/100
The gist: The paper addresses a limitation in On-Policy Self-Distillation (OPSD) for multimodal large language models (MLLMs).
Terminology
Summary
The paper addresses a limitation in On-Policy Self-Distillation (OPSD) for multimodal large language models (MLLMs). OPSD has become a standard post-training approach: it samples trajectories from the current student and uses a copy of the same model to provide dense targets at student-visited prefixes.
While existing methods construct privileged information from diverse sources (e.g., verified textual solutions or transformed visual inputs such as evidence-centered crops in Vision-OPD or privileged visual thoughts in Visual-OPSD), the authors argue these designs overlook Modality Imbalance,
a challenge inherent to MLLM reasoning.
They define Modality Imbalance as: the model relies disproportionately on textual context instead of integrating the multimodal input.
Because textual priors can dominate generation even when the task depends on the image, "carefully designed privileged information can enrich the teacher input without revealing how the visual and textual modalities contribute to each prediction. Consequently, the model can underuse this information, limiting the effectiveness of OPSD." The central question is: How can privileged information be designed to remain effective for MLLM OPSD under Modality Imbalance?
The authors construct a Positive Teacher with the Zoom-In Image and a Negative Teacher with the Mask Image to expose different degrees of Modality Imbalance. They define the Modality-Balance Attention Ratio as visual attention relative to textual attention over an on-policy trajectory.
Their empirical analysis (Figure 1) shows: "the ratio is lowest for the Negative Teacher, intermediate for the student, and highest for the Positive Teacher. This ordering shows that the matched image conditions expose distinct degrees of Modality Imbalance through explicit differences in the model input." Moreover, the Modality-Balance Logits Margin—measuring how differently the Positive and Negative Teachers score the student response—correlates with student correctness: "From low to high margin intervals, the gap between their Modality-Balance Attention Ratios widens, while student correctness increases. Teachers with different levels of Modality Balance thus provide a directional scoring distinction that can serve as privileged information for OPSD."
OPD-V instantiates Modality Balance as privileged information. The architecture: "The student receives the Original Image, the Positive Teacher receives the Zoom-In Image, and the Negative Teacher receives the Mask Image. All three distributions score the same student-generated tokens under matched textual context and on-policy prefixes."
Positive/Negative Teachers: The Zoom-In Image magnifies the task-relevant region; the Mask Image is obtained by replacing a random rectangular region in the Zoom-In Image with black pixels,
creating a visually weakened condition while preserving most image context.
Modality-Balance Trust Region: The tokenwise Modality-Balance Logits Margin is computed as δ t MB = log q t+(y t) − log q t−(y t). "A positive Logits Margin indicates that the same token receives more support from the Positive Teacher than from the Negative Teacher. These positions form the Modality-Balance Trust Region: R MB(y) = t ∈ T y δ t MB > 0."
Objective: Within the trust region, OPD-V applies Jensen–Shannon distillation from the Positive Teacher to the student, scaled by the margin:
L OPD-V(θ) = E[1/Σ r t · Σ t∈R MB(y) r t δ t MB D JS(q t+, p θ,t)]
The loss uses top-K (K=100) tail-adjusted distillation for memory efficiency. The detached teacher is updated via EMA (rate 0.05). Both teachers share the same EMA parameters; OPD-V therefore introduces no additional model parameters or student updates.
-
We identify Modality Imbalance as a limitation that restricts the effectiveness of privileged information in MLLM OPSD.
-
We show that Modality Balance can serve as privileged information and propose OPD-V, which instantiates it through the Positive Teacher, Negative Teacher, and Modality-Balance Trust Region.
-
We demonstrate consistent improvements across representative OPSD methods while reducing training cost.
OPD-V is evaluated across 6 benchmarks (V* Bench, ZoomBench, HR-Bench 4K/8K, MME-RealWorld EN/CN), 4 MLLM backbones (Qwen3.5-4B, Qwen3.5-9B, Qwen3-VL-4B-Instruct, Qwen3-VL-8B-Instruct), and 5 post-training methods (SFT, GRPO, OPSD, Vision-OPD, VA-OPD).
Key results:
-
Qwen3.5-4B: average accuracy rises
from 64.30% to 80.01%, an absolute gain of 15.71 percentage points,
exceeding the strongest alternative baseline (Vision-OPD at 77.10%) by 2.91 points. The 4B modeloutperforming every closed-source model listed in Table 1
and surpassingsubstantially larger open-source systems, including the 397B Qwen3.5 model at 77.44% and the 1T-parameter Kimi-K2.6 model at 72.81%.
-
Other backbones:
OPD-V improves Qwen3-VL-4B from 68.00% to 74.14%, Qwen3-VL-8B from 68.93% to 73.03%, and Qwen3.5-9B from 69.75% to 77.63%.
-
Training dynamics: OPD-V maintains concise responses (
mean response length is 140.9 tokens for OPD-V and 553.8 tokens for OPSD, a 74.5% reduction
), stable policy entropy (rolling means 0.827 for 4B and 0.761 for 9B), and a persistent trust region (fraction averages 53.8% for 4B, 49.6% for 9B).
Mean step latency decreases from 352 s to 240 s on Qwen3.5-4B and from 451 s to 340 s on Qwen3.5-9B, corresponding to reductions of 31.8% and 24.7%.
This is due to shorter rollouts and more efficient privileged-input construction: "teacher-batch construction and preprocessing decrease from 79.2 s for OPSD to 15.3 s for OPD-V. Although OPD-V evaluates two teachers, their combined forward time is 27.1 s, below the 40.4 s required by the single OPSD teacher."
Ablations show teacher complementarity: The Negative Teacher alone improves the base model from 64.30% to 71.98%, while the Positive Teacher alone reaches 74.62%. Combining them yields 80.01%.
The Zoom-In + Mask operation pair is strongest: replacing Zoom-In with Repeat Image drops accuracy to 75.95%; replacing Mask with Blur (74.31%), Prune (73.74%), or No Image (72.17%) all underperform.
The paper concludes: "This work identifies Modality Balance as privileged information for OPSD and introduces OPD-V, which uses the Positive Teacher and Negative Teacher to define a Modality-Balance Trust Region. Across six benchmarks and four MLLM backbones, OPD-V consistently improves reasoning performance while reducing training cost."
Improvements for AI systems
By incorporating OPD-V into a multimodal LLM training pipeline, I can make the following concrete improvements:
- Add a Positive/Negative Teacher pair for privileged visual supervision.
-
The Positive Teacher sees a zoomed-in crop of the task-relevant image region.
-
The Negative Teacher sees the same zoomed-in image with a random rectangular region masked out.
-
Both teachers score the student’s own generated tokens under identical textual prefixes, so the student learns from the difference between visually strong and visually weakened conditions.
- Construct a Modality-Balance Trust Region dynamically during training.
-
For each token position, compute the logits margin:
log p positive(y t) − log p negative(y t). -
Only distill tokens where the margin is positive, i.e., where the visual signal actually supports the token.
-
This prevents the student from copying text-prior-dominated answers and forces it to rely on visual evidence.
- Scale distillation by the Modality-Balance margin.
-
Use the margin as a token-level weight in the Jensen–Shannon distillation loss.
-
Tokens with stronger visual support receive stronger supervision; ambiguous tokens are down-weighted.
-
This yields a stable, self-regularizing training signal that resists the usual text-prior collapse.
- Reduce training cost while improving visual reasoning.
-
Teacher-batch construction is far cheaper: no need for expensive privileged-image preprocessing pipelines.
-
The method produces much shorter rollouts (e.g., 141 tokens vs 554 tokens), cutting step latency by 25–32%.
-
I can train a 4B MLLM to outperform far larger models on fine-grained visual benchmarks (e.g., >80% average accuracy on V* Bench, ZoomBench, and HR-Bench family).
- Maintain response quality and policy stability.
-
Use EMA-shared teachers (rate 0.05) with no extra trainable parameters.
-
Track policy entropy to avoid collapse while keeping the trust region active across training (roughly 50% of tokens on average).
-
The result is a model that produces concise, grounded answers instead of verbose text-prior explanations.
The improved AI system can:
-
Correctly answer fine-grained visual questions that depend on small image regions, not just global context.
-
Refuse or down-weight text-only shortcuts when the image is ambiguous or visually weakened.
-
Generate shorter, more direct responses with higher visual-grounded accuracy.
-
Be trained faster and more cheaply than standard OPSD, while matching or surpassing much larger closed- and open-source models.
-
Generalize across multiple backbones (4B/8B/9B) and multiple post-training methods (SFT, GRPO, OPSD, Vision-OPD, VA-OPD), making it a drop-in improvement for existing MLLM training systems.
Sources
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
- Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning
- Visual-Advantage On-Policy Distillation for Vision-Language Models
- Visual Contrastive Self-Distillation
- Evaluating and Steering Modality Preferences in Multimodal Large Language Model
- Qwen3-VL Technical Report
- Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
- Thyme: Think Beyond Images
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- DeepEyesV2: Toward Agentic Multimodal Model
- MiMo-VL Technical Report
- SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- Distilling the Knowledge in a Neural Network
- CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models