Is Better Teacher Supervision Enough? Unlocking Student-side Learning in Multimodal On-Policy Distillation
cs.CV
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/Sirilaw/S-OPD
Terminology
Sources
- OPD-V: Visual On-Policy Self-Distillation with Modality Balance
- Qwen3-VL Technical Report
- Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Distilling the Knowledge in a Neural Network
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Visual Contrastive Self-Distillation
- Visual-Advantage On-Policy Distillation for Vision-Language Models
- Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
- ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
- V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning
- What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation
- ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
- Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
- LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models