OPD-V: Visual On-Policy Self-Distillation with Modality Balance

arXiv:2608.05131 · cs.CV, cs.AI · Submitted 2026-08-05 · Read on arXiv

Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua

National University of Singapore · Ludwig Maximilian University of Munich · Munich Center for Machine Learning · Sun Yat-sen University

cs.CV, cs.AI

Submitted: 2026-08-05

Comments: Corrected the uploaded manuscript. Project Page:https://github.com/aniri15/OPD-V

Code: https://github.com/aniri15/OPD-V

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 80/100

The gist: The paper addresses a limitation in On-Policy Self-Distillation (OPSD) for multimodal large language models (MLLMs).

Terminology

Summary

The paper addresses a limitation in On-Policy Self-Distillation (OPSD) for multimodal large language models (MLLMs). OPSD has become a standard post-training approach: it samples trajectories from the current student and uses a copy of the same model to provide dense targets at student-visited prefixes. While existing methods construct privileged information from diverse sources (e.g., verified textual solutions or transformed visual inputs such as evidence-centered crops in Vision-OPD or privileged visual thoughts in Visual-OPSD), the authors argue these designs overlook Modality Imbalance, a challenge inherent to MLLM reasoning.

They define Modality Imbalance as: the model relies disproportionately on textual context instead of integrating the multimodal input. Because textual priors can dominate generation even when the task depends on the image, "carefully designed privileged information can enrich the teacher input without revealing how the visual and textual modalities contribute to each prediction. Consequently, the model can underuse this information, limiting the effectiveness of OPSD." The central question is: How can privileged information be designed to remain effective for MLLM OPSD under Modality Imbalance?

The authors construct a Positive Teacher with the Zoom-In Image and a Negative Teacher with the Mask Image to expose different degrees of Modality Imbalance. They define the Modality-Balance Attention Ratio as visual attention relative to textual attention over an on-policy trajectory.

Their empirical analysis (Figure 1) shows: "the ratio is lowest for the Negative Teacher, intermediate for the student, and highest for the Positive Teacher. This ordering shows that the matched image conditions expose distinct degrees of Modality Imbalance through explicit differences in the model input." Moreover, the Modality-Balance Logits Margin—measuring how differently the Positive and Negative Teachers score the student response—correlates with student correctness: "From low to high margin intervals, the gap between their Modality-Balance Attention Ratios widens, while student correctness increases. Teachers with different levels of Modality Balance thus provide a directional scoring distinction that can serve as privileged information for OPSD."

OPD-V instantiates Modality Balance as privileged information. The architecture: "The student receives the Original Image, the Positive Teacher receives the Zoom-In Image, and the Negative Teacher receives the Mask Image. All three distributions score the same student-generated tokens under matched textual context and on-policy prefixes."

Positive/Negative Teachers: The Zoom-In Image magnifies the task-relevant region; the Mask Image is obtained by replacing a random rectangular region in the Zoom-In Image with black pixels, creating a visually weakened condition while preserving most image context.

Modality-Balance Trust Region: The tokenwise Modality-Balance Logits Margin is computed as δ t MB = log q t+(y t) − log q t−(y t). "A positive Logits Margin indicates that the same token receives more support from the Positive Teacher than from the Negative Teacher. These positions form the Modality-Balance Trust Region: R MB(y) = t ∈ T y δ t MB > 0."

Objective: Within the trust region, OPD-V applies Jensen–Shannon distillation from the Positive Teacher to the student, scaled by the margin:

L OPD-V(θ) = E[1/Σ r t · Σ t∈R MB(y) r t δ t MB D JS(q t+, p θ,t)]

The loss uses top-K (K=100) tail-adjusted distillation for memory efficiency. The detached teacher is updated via EMA (rate 0.05). Both teachers share the same EMA parameters; OPD-V therefore introduces no additional model parameters or student updates.

  1. We identify Modality Imbalance as a limitation that restricts the effectiveness of privileged information in MLLM OPSD.

  2. We show that Modality Balance can serve as privileged information and propose OPD-V, which instantiates it through the Positive Teacher, Negative Teacher, and Modality-Balance Trust Region.

  3. We demonstrate consistent improvements across representative OPSD methods while reducing training cost.

OPD-V is evaluated across 6 benchmarks (V* Bench, ZoomBench, HR-Bench 4K/8K, MME-RealWorld EN/CN), 4 MLLM backbones (Qwen3.5-4B, Qwen3.5-9B, Qwen3-VL-4B-Instruct, Qwen3-VL-8B-Instruct), and 5 post-training methods (SFT, GRPO, OPSD, Vision-OPD, VA-OPD).

Key results:

  • Qwen3.5-4B: average accuracy rises from 64.30% to 80.01%, an absolute gain of 15.71 percentage points, exceeding the strongest alternative baseline (Vision-OPD at 77.10%) by 2.91 points. The 4B model outperforming every closed-source model listed in Table 1 and surpassing substantially larger open-source systems, including the 397B Qwen3.5 model at 77.44% and the 1T-parameter Kimi-K2.6 model at 72.81%.

  • Other backbones: OPD-V improves Qwen3-VL-4B from 68.00% to 74.14%, Qwen3-VL-8B from 68.93% to 73.03%, and Qwen3.5-9B from 69.75% to 77.63%.

  • Training dynamics: OPD-V maintains concise responses (mean response length is 140.9 tokens for OPD-V and 553.8 tokens for OPSD, a 74.5% reduction), stable policy entropy (rolling means 0.827 for 4B and 0.761 for 9B), and a persistent trust region (fraction averages 53.8% for 4B, 49.6% for 9B).

Mean step latency decreases from 352 s to 240 s on Qwen3.5-4B and from 451 s to 340 s on Qwen3.5-9B, corresponding to reductions of 31.8% and 24.7%. This is due to shorter rollouts and more efficient privileged-input construction: "teacher-batch construction and preprocessing decrease from 79.2 s for OPSD to 15.3 s for OPD-V. Although OPD-V evaluates two teachers, their combined forward time is 27.1 s, below the 40.4 s required by the single OPSD teacher."

Ablations show teacher complementarity: The Negative Teacher alone improves the base model from 64.30% to 71.98%, while the Positive Teacher alone reaches 74.62%. Combining them yields 80.01%. The Zoom-In + Mask operation pair is strongest: replacing Zoom-In with Repeat Image drops accuracy to 75.95%; replacing Mask with Blur (74.31%), Prune (73.74%), or No Image (72.17%) all underperform.

The paper concludes: "This work identifies Modality Balance as privileged information for OPSD and introduces OPD-V, which uses the Positive Teacher and Negative Teacher to define a Modality-Balance Trust Region. Across six benchmarks and four MLLM backbones, OPD-V consistently improves reasoning performance while reducing training cost."

Improvements for AI systems

By incorporating OPD-V into a multimodal LLM training pipeline, I can make the following concrete improvements:

  1. Add a Positive/Negative Teacher pair for privileged visual supervision.
  • The Positive Teacher sees a zoomed-in crop of the task-relevant image region.

  • The Negative Teacher sees the same zoomed-in image with a random rectangular region masked out.

  • Both teachers score the student’s own generated tokens under identical textual prefixes, so the student learns from the difference between visually strong and visually weakened conditions.

  1. Construct a Modality-Balance Trust Region dynamically during training.
  • For each token position, compute the logits margin: log p positive(y t) − log p negative(y t).

  • Only distill tokens where the margin is positive, i.e., where the visual signal actually supports the token.

  • This prevents the student from copying text-prior-dominated answers and forces it to rely on visual evidence.

  1. Scale distillation by the Modality-Balance margin.
  • Use the margin as a token-level weight in the Jensen–Shannon distillation loss.

  • Tokens with stronger visual support receive stronger supervision; ambiguous tokens are down-weighted.

  • This yields a stable, self-regularizing training signal that resists the usual text-prior collapse.

  1. Reduce training cost while improving visual reasoning.
  • Teacher-batch construction is far cheaper: no need for expensive privileged-image preprocessing pipelines.

  • The method produces much shorter rollouts (e.g., 141 tokens vs 554 tokens), cutting step latency by 25–32%.

  • I can train a 4B MLLM to outperform far larger models on fine-grained visual benchmarks (e.g., >80% average accuracy on V* Bench, ZoomBench, and HR-Bench family).

  1. Maintain response quality and policy stability.
  • Use EMA-shared teachers (rate 0.05) with no extra trainable parameters.

  • Track policy entropy to avoid collapse while keeping the trust region active across training (roughly 50% of tokens on average).

  • The result is a model that produces concise, grounded answers instead of verbose text-prior explanations.

The improved AI system can:

  • Correctly answer fine-grained visual questions that depend on small image regions, not just global context.

  • Refuse or down-weight text-only shortcuts when the image is ambiguous or visually weakened.

  • Generate shorter, more direct responses with higher visual-grounded accuracy.

  • Be trained faster and more cheaply than standard OPSD, while matching or surpassing much larger closed- and open-source models.

  • Generalize across multiple backbones (4B/8B/9B) and multiple post-training methods (SFT, GRPO, OPSD, Vision-OPD, VA-OPD), making it a drop-in improvement for existing MLLM training systems.

Sources

Related papers