VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
Kangning Zhang, Yixing Li, Shuai Shao, Qingyao Li, Zhengxi Lu, Zhiyuan Yao, Jianghao Lin, Wenxiang Jiao, Yuan Lu, Weiwen Liu, Weinan Zhang, Yong Yu
cs.CV, cs.CL
Submitted: 2026-07-30
Comments: The project is accessible at https://github.com/DeepExperience/VAD_Multimodal_OPD
Code: https://github.com/DeepExperience/VAD_Multimodal_OPD
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher.
Terminology
Abstract
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning
- ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
- Thinking with Deltas: Incentivizing Reinforcement Learning via Differential Visual Reasoning Policy
- DeepEyesV2: Toward Agentic Multimodal Model
- Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
- Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning
- Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
- Visual Instruction Tuning
- Visual-Advantage On-Policy Distillation for Vision-Language Models
- gpt-oss-120b & gpt-oss-20b Model Card
- Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation
- Contrastive On-Policy Distillation
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning
- ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models