Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models
cs.CV, cs.AI
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- Qwen2.5-VL Technical Report
- SAM 3: Segment Anything with Concepts
- From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
- Training Large Language Models to Reason in a Continuous Latent Space
- MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation
- Latent Visual Reasoning
- GRES: Generalized Referring Expression Segmentation
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
- RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
- Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
- AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation
- GLaMM: Pixel Grounding Large Multimodal Model
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- LLM-Seg: Bridging Image Segmentation and Large Language Model Reasoning
- PixelThink: Towards Efficient Chain-of-Pixel Reasoning
- SegLLM: Multi-round Reasoning Segmentation
- InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
- See, Say, and Segment: Teaching LMMs to Overcome False Premises
- GSVA: Generalized Segmentation via Multimodal Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models