Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs
cs.CL
Submitted: 2026-04-08
Updated: 2026-08-30
Terminology
Sources
- Latent Visual Reasoning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Qwen3-VL Technical Report
- ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
- Monet: Reasoning in Latent Visual Space Beyond Images and Language
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
- Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
- Multimodal Chain-of-Thought Reasoning in Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering