WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- Qwen2.5-VL Technical Report
- Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
- ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
- Thinking with Generated Images
- Emerging Properties in Unified Multimodal Pretraining
- Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models
- Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process
- DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models
- LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
- Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
- BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
- Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
- What's Holding Back Latent Visual Reasoning?
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs
- When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
- Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs
- Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models