On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
cs.CL, cs.CV
Submitted: 2026-05-28
Updated: 2026-09-02
Comments: Project: https://asymmetric-vlm-post-training.github.io/
Project page: https://asymmetric-vlm-post-training.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Post-training has greatly improved reasoning in frontier vision-language models, yet its gains for perception remain comparatively limited, creating a bottleneck for end-to-end visual reasoning.
Terminology
Abstract
Post-training has greatly improved reasoning in frontier vision-language models, yet its gains for perception remain comparatively limited, creating a bottleneck for end-to-end visual reasoning. To investigate this gap, we introduce a controlled diagnostic framework with two synthetic tasks that disentangle perception from reasoning. Our analysis reveals a consistent perception-reasoning asymmetry: post-training improves reasoning more substantially than perception, though the underlying mechanism differs across training paradigms. For supervised fine-tuning (SFT), this asymmetry stems from token imbalance, with perception occupying a smaller fraction of tokens in chain-of-thought supervision. Reweighting the loss boosts end-to-end performance by up to 18.2 points. For reinforcement learning (RL), the asymmetry instead arises from reward coupling, as outcome rewards correlate more strongly with reasoning than perception. Adding a perception-aware reward improves end-to-end accuracy by up to 6.0 points; when ground-truth perception rewards are unavailable, a reliable surrogate provides useful signal, yielding gains of 2.2 points. Beyond the controlled setting, these strategies also improve real-world visual reasoning, with gains of up to 3.3 points across three benchmarks. Overall, we diagnose the causes of asymmetric optimization and provide actionable guidance that benefits both synthetic and realistic settings.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Hallucination of Multimodal Large Language Models: A Survey
- Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- OpenAI o1 System Card
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model
- What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis
- Unleashing Perception-Time Scaling to Multimodal Reasoning Models
- Self-Rewarding Vision-Language Model via Reasoning Decomposition
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
- A Survey on Hallucination in Large Vision-Language Models
- MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
- LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering