MM-OPD: Towards One More Bottleneck Between Perception and Reasoning
cs.CV, cs.AI
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Reinforcement Learning via Self-Distillation
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- V-Thinker: Interactive Thinking with Images
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Purified OPSD: On-Policy Self-Distillation Without Losing How to Think
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
- Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
- Self-Distilled RLVR
- DOPD: Dual On-policy Distillation
- Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- Group Sequence Policy Optimization
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models