Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning
cs.LG, cs.CL
Submitted: 2026-09-17
Updated: 2026-09-29
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multimodal reasoning requires models to draw on information from multiple modalities throughout the reasoning process.
Terminology
Abstract
Multimodal reasoning requires models to draw on information from multiple modalities throughout the reasoning process. Yet existing methods often concatenate modality-specific thought tokens in a single sequence, leaving the model to bridge representational differences as it reasons across modalities. We introduce Uni-LaDiR (Unified Latent Diffusion Reasoner), a framework that brings these thoughts into a shared latent space for reasoning. A unified encoder maps teacher reasoning steps from different modalities into shared thought tokens, trained to preserve the information needed for later reasoning steps and the final answer or action. Because the same context can support multiple valid next steps, we use diffusion to predict the next block of thought tokens from the input and preceding blocks. Jointly training the encoder and diffusion reasoner with shared model weights encourages thought tokens to be both useful for the task and predictable from the available context. At inference, the model generates these tokens without teacher observations. Across eleven vision-language model (VLM) benchmarks and two vision-language-action (VLA) suites, Uni-LaDiR achieves relative gains over the strongest evaluated baselines of 7.3% on visual reasoning tasks and 6.1% on robot manipulation tasks.
Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
- Qwen3-VL Technical Report
- Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
- Large Concept Models: Language Modeling in a Sentence Representation Space
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- PaLI: A Jointly-Scaled Multilingual Language-Image Model
- PaLI-X: On Scaling up a Multilingual Vision and Language Model
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Universal Transformers
- Emerging Properties in Unified Multimodal Pretraining
- LLM Latent Reasoning as Chain of Superposition
- Implicit Chain of Thought Reasoning via Knowledge Distillation
- From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
- PaLM-E: An Embodied Multimodal Language Model
- End-to-End Training for Unified Tokenization and Latent Denoising
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks