Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models
summary
The gist
Vision-language models (VLMs) excel at many tasks, yet continue to struggle with spatial reasoning—problems where the key information is not directly observable in the input.
In short
The episode discusses the paper 'Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models.' It addresses how AI models struggle with spatial reasoning—predictive tasks like seeing a room from a new angle. The hosts conclude that training models to generate internal images (perception tokens) allows them to build better mental maps of space, significantly boosting accuracy over traditional text-based reasoning.
Key concepts
- Spatial Reasoning Difficulty
- Current AI models excel at object recognition but fail when asked how a scene changes based on movement or perspective. They lack the ability to build a complete mental model of a space that goes beyond what is immediately visible in the input image.
- Imaginative Perception Tokens
- This is a training technique where the the model generates an internal image (a 'perception token') of what it believes it would see from a new viewpoint. This forces the AI to construct a mental scene before answering, allowing it to better understand complex spatial relationships.
- Text Chain-of-Thought
- This method requires the model to explain its reasoning using words before providing an answer. The hosts noted that for geometric tasks, this approach often hurts performance because describing a viewpoint rotation in words makes it difficult for models to build the necessary mental picture.
Terminology used across episodes
This episode discusses
- Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models · Paper Radio
- Emerging Properties in Unified Multimodal Pretraining
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
- AI2-THOR: An Interactive 3D Environment for Visual AI
- Qwen3-VL Technical Report
- Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
- Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts · Paper Radio
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- Show-o2: Improved Native Unified Multimodal Models
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- Visual Spatial Tuning
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
- MindCube: Spatial Mental Modeling from Limited Views
- Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
- ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
- Visual Spatial Reasoning
The paper
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models · Read on arXiv
Mahtab Bigverdi, Linjie Li, Weikai Huang, Yiming Liu, Jaemin Cho, Tuhin Kundu, Chris Dongjoo Kim, Zelun Luo, Jieyu Zhang, Linda Shapiro, Ranjay Krishna
University of Washington · Allen Institute for AI · Microsoft · OpenAI
Vision language models (VLMs) excel at many tasks but still struggle with spatial reasoning when critical information is not directly observable. Many such problems require imaginative perception: inferring what would be seen from an unseen viewpoint, tracing paths through occluded spaces, or integrating partial observations into a coherent spatial representation. We introduce Imaginative Perception Tokens (IPT), intermediate perceptual representations that externalize what a VLM would perceive under alternative spatial configurations while remaining consistent with the observed input. To study this capability, we formulate three tasks, Perspective Taking (PET), Path Tracing (PT), and Multiview Counting (MVC), and construct datasets of approximately 20K examples with ground truth imaginations, answers, and evaluation benchmarks. Using the unified VLM BAGEL as the backbone, IPT supervision consistently improves spatial reasoning and often outperforms textual chain of thought training, even without generating images at inference time. On MVC, IPT improves accuracy by 3.4% and achieves competitive performance with strong closed-source models on PT. We further find that combining IPT and label-only supervision yields additional gains, whereas textual chain of thought can substantially degrade performance, suggesting a modality mismatch when spatial computation is forced through language. Overall, IPT provides a principled supervision signal for reasoning about unobserved spatial structure, improving generalization while producing interpretable intermediate representations.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models".
Jane: The paper was written by Mahtab Bigverdi, Linjie Li, Weikai Huang, Yiming Liu, Jaemin Cho et al. from University of Washington and Allen Institute for AI and Microsoft and OpenAI.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a title that really makes you stop and think: "Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models." Jane, I have to say, just reading that title got me excited.
Jane: It got me excited too, Tom. And I think the title is actually perfect because it tells you exactly what the problem is. These AI models, they're great at recognizing objects in a picture, but if you ask them something like, "If I walk over there and turn around, what will I see?" they just fall apart. The paper is saying we need to give them imagination.
Tom: Right, and that's the key word, imagination. But let's be specific here. The paper is from a big team at the University of Washington, the Allen Institute, Microsoft, and OpenAI. They're not just theorizing about this. They built three new tasks to test this exact ability.
Jane: And those tasks are what really sold me on this. There's Perspective Taking, where you see a room and you have to figure out what happens when you move to a marked spot and turn. There's Path Tracing, where you're looking at a top-down map and you have to figure out what you'd see from ground level at a specific point. And there's Multiview Counting, where you see four different photos and you have to figure out how many chairs are in the whole room, even though each photo only shows part of it.
Tom: So it's not just about recognizing what's in front of you. It's about building a mental model of the space that goes beyond what you can actually see. That's the "imaginative" part of the title.
Jane: Exactly. And the "tokens" part is about how they train the model to do this. They don't just ask for the answer. They ask the model to first generate an image of what it thinks it would see from that new position. That image is the "perception token." It's a way of forcing the model to actually construct the scene in its mind before answering.
Tom: And the results, we're going to get into those in a second, but the short version is that this approach works a lot better than just asking the model to think in text. It's a really clever idea, and I can't wait to dig into the details.
Jane: Me neither. Let's get into the actual summary of what they found.
Summary: Jane: So, Tom, we've talked about the title and the core idea. Now let's talk about what actually happened when they ran the experiments. The results are in the paper, and they're pretty striking.
Tom: They are. And I think the most important number to start with is how badly the current models do. They tested GPT-five Gemini, Qwen, all the big names. And on these tasks, they're barely above chance. On Perspective Taking, most of them are hovering around fifty percent, which is literally a coin flip.
Jane: That's the baseline, and it's important because it shows these tasks are genuinely hard. They're not testing something trivial. But then they take their own model, which is based on something called BAGEL, and they fine-tune it. And the jump is enormous. On Perspective Taking, they go from about forty percent accuracy to ninety-seven point five percent just by training on the answers.
Lu: If I can jump in here, Jane, that jump is the first big story. It shows that the capability isn't missing entirely, it's just dormant. The model can learn it if you give it the right data. But the second part is what's really interesting to me. When they add the imaginative perception tokens, the intermediate images, they get another boost on the harder tasks.
Tom: Right, and that boost is the key. On Multiview Counting, the image-based training gets them to sixty-seven point three percent, which beats the answer-only training at sixty-three point nine percent. And on Path Tracing, the mixed approach gets them to sixty-six point seven percent. But here's the wild part, Jane.
Jane: What's that, Tom?
Tom: The model doesn't even need to generate the image at test time to get the benefit. When they train it to generate the image but then at test time just ask for the answer directly, it still performs almost as well. The training with images makes the model's internal representations better, even if it never shows its work.
Lu: That's the part that really excites me, Tom. It suggests that the imagination isn't just a trick. It's actually reshaping how the model thinks about space internally. The image generation is a training tool that builds a better mental model, not just a display mechanism.
Jane: And that's a huge deal. It means you get the benefit of the visual reasoning without the computational cost of actually generating an image every time you ask a question. That's a win for both accuracy and efficiency.
Tom: So we've got the problem, we've got the solution, and we've got the numbers. But I want to talk about what this means for how we think about AI reasoning in general. Let's get into the improvements and the bigger picture.
Improvements: Tom: So we've established that this works. The imaginative perception tokens give a real boost. But I want to dig into why it works, because that's where the real insight is. Jane, what do you make of the comparison they did with text-based reasoning?
Jane: Oh, that's the part that really made me sit up. They compared their image-based approach against something called "text chain-of-thought." That's where you ask the model to explain its reasoning in words before giving the answer. And the text approach actually hurt performance.
Tom: It hurt it a lot. On Perspective Taking, the text reasoning dropped them from ninety-seven point five percent down to eighty-three point one percent. That's a massive drop. And on Path Tracing, it was even worse, down to forty-nine point seven percent, which is basically random guessing.
Lu: And I think the reason is pretty clear. Spatial reasoning is fundamentally geometric. When you try to describe a viewpoint rotation in words, you lose the spatial relationships. It's like trying to describe a spiral staircase to someone who's never seen one. You can say the words, but they can't build the image in their head. The image token bypasses that problem entirely.
Jane: Exactly. And the paper actually has a great way of showing this. They gave the model the ground-truth image, the perfect imagination, and then asked it to answer. On Path Tracing, accuracy jumped to eighty-six point seven percent. That's a huge jump from the fifty percent when the model had to generate its own image.
Tom: So that tells us the bottleneck isn't the reasoning, it's the imagination quality. If the model can imagine the scene correctly, it can answer correctly. The problem is that generating that image is hard.
Meng: And that's where I want to ask a practical question, if I can. You've got this great result, but what does it cost? Generating images is computationally expensive. Is this something that can actually run in a real product, or is it just a research curiosity?
Tom: That's a fair question, Meng. And the paper has a good answer. They found that you don't need to generate the image at test time to get the benefit. The model trained with images does better even when it just answers directly. So you can train with the expensive method and then deploy a model that's fast at inference.
Meng: That's a relief. So the training cost is higher, but the deployment cost is the same as a standard model, just with better accuracy. That makes it practical.
Lu: And it also opens up a really interesting research direction. If we can get the imagination quality higher, the accuracy should go up too. The paper shows there's still a big gap between what the model imagines and the ground truth. Closing that gap is the next big challenge.
Jane: So we've got a method that works, we know why it works, and we know where the remaining challenges are. Let's wrap this up and think about what it all means.
Conclusion: Jane: Alright, Tom, let's bring it home. We've spent the show talking about "Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models," and I think we've only scratched the surface.
Tom: We really have. And I think the core message is simple. If you want AI to reason about space, you need to let it think in space. Text is a terrible medium for describing geometry. Images are the natural language of spatial reasoning.
Jane: And the paper shows this isn't just a nice idea. It's a measurable improvement. They beat the best closed-source models on Path Tracing. They show that training with imagination transfers to other tasks they weren't even trained on. This is a real step forward.
Lu: And I think the long-term impact is even bigger. This isn't just about counting chairs or navigating rooms. It's about building models that can construct mental models of the world. That's a foundational capability for robotics, for augmented reality, for any system that needs to understand and interact with physical space.
Meng: And from my side, the fact that you can get the benefit without the inference cost makes this something that could actually ship. It's not just a lab experiment. It's a practical technique.
Tom: So we're saying goodbye to this paper, but the ideas in it are going to stick with us. The next time you see a robot that can navigate a room it's never seen, or an AR system that can predict what's around a corner, you'll know where the seed was planted.
Jane: And that's the exciting part. This paper isn't the end of the story. It's the beginning of a new way to think about how AI understands space. Thanks for joining us, everyone. We'll see you next time with another paper to break down.
Tom: Take care, folks.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language