Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models

arXiv:2606.03988 · cs.AI · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models".

Jane: The paper was written by Mahtab Bigverdi, Linjie Li, Weikai Huang, Yiming Liu, Jaemin Cho et al. from University of Washington and Allen Institute for AI and Microsoft and OpenAI.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a title that really makes you stop and think: "Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models." Jane, I have to say, just reading that title got me excited.

Jane: It got me excited too, Tom. And I think the title is actually perfect because it tells you exactly what the problem is. These AI models, they're great at recognizing objects in a picture, but if you ask them something like, "If I walk over there and turn around, what will I see?" they just fall apart. The paper is saying we need to give them imagination.

Tom: Right, and that's the key word, imagination. But let's be specific here. The paper is from a big team at the University of Washington, the Allen Institute, Microsoft, and OpenAI. They're not just theorizing about this. They built three new tasks to test this exact ability.

Jane: And those tasks are what really sold me on this. There's Perspective Taking, where you see a room and you have to figure out what happens when you move to a marked spot and turn. There's Path Tracing, where you're looking at a top-down map and you have to figure out what you'd see from ground level at a specific point. And there's Multiview Counting, where you see four different photos and you have to figure out how many chairs are in the whole room, even though each photo only shows part of it.

Tom: So it's not just about recognizing what's in front of you. It's about building a mental model of the space that goes beyond what you can actually see. That's the "imaginative" part of the title.

Jane: Exactly. And the "tokens" part is about how they train the model to do this. They don't just ask for the answer. They ask the model to first generate an image of what it thinks it would see from that new position. That image is the "perception token." It's a way of forcing the model to actually construct the scene in its mind before answering.

Tom: And the results, we're going to get into those in a second, but the short version is that this approach works a lot better than just asking the model to think in text. It's a really clever idea, and I can't wait to dig into the details.

Jane: Me neither. Let's get into the actual summary of what they found.

Summary: Jane: So, Tom, we've talked about the title and the core idea. Now let's talk about what actually happened when they ran the experiments. The results are in the paper, and they're pretty striking.

Tom: They are. And I think the most important number to start with is how badly the current models do. They tested GPT-five Gemini, Qwen, all the big names. And on these tasks, they're barely above chance. On Perspective Taking, most of them are hovering around fifty percent, which is literally a coin flip.

Jane: That's the baseline, and it's important because it shows these tasks are genuinely hard. They're not testing something trivial. But then they take their own model, which is based on something called BAGEL, and they fine-tune it. And the jump is enormous. On Perspective Taking, they go from about forty percent accuracy to ninety-seven point five percent just by training on the answers.

Lu: If I can jump in here, Jane, that jump is the first big story. It shows that the capability isn't missing entirely, it's just dormant. The model can learn it if you give it the right data. But the second part is what's really interesting to me. When they add the imaginative perception tokens, the intermediate images, they get another boost on the harder tasks.

Tom: Right, and that boost is the key. On Multiview Counting, the image-based training gets them to sixty-seven point three percent, which beats the answer-only training at sixty-three point nine percent. And on Path Tracing, the mixed approach gets them to sixty-six point seven percent. But here's the wild part, Jane.

Jane: What's that, Tom?

Tom: The model doesn't even need to generate the image at test time to get the benefit. When they train it to generate the image but then at test time just ask for the answer directly, it still performs almost as well. The training with images makes the model's internal representations better, even if it never shows its work.

Lu: That's the part that really excites me, Tom. It suggests that the imagination isn't just a trick. It's actually reshaping how the model thinks about space internally. The image generation is a training tool that builds a better mental model, not just a display mechanism.

Jane: And that's a huge deal. It means you get the benefit of the visual reasoning without the computational cost of actually generating an image every time you ask a question. That's a win for both accuracy and efficiency.

Tom: So we've got the problem, we've got the solution, and we've got the numbers. But I want to talk about what this means for how we think about AI reasoning in general. Let's get into the improvements and the bigger picture.

Improvements: Tom: So we've established that this works. The imaginative perception tokens give a real boost. But I want to dig into why it works, because that's where the real insight is. Jane, what do you make of the comparison they did with text-based reasoning?

Jane: Oh, that's the part that really made me sit up. They compared their image-based approach against something called "text chain-of-thought." That's where you ask the model to explain its reasoning in words before giving the answer. And the text approach actually hurt performance.

Tom: It hurt it a lot. On Perspective Taking, the text reasoning dropped them from ninety-seven point five percent down to eighty-three point one percent. That's a massive drop. And on Path Tracing, it was even worse, down to forty-nine point seven percent, which is basically random guessing.

Lu: And I think the reason is pretty clear. Spatial reasoning is fundamentally geometric. When you try to describe a viewpoint rotation in words, you lose the spatial relationships. It's like trying to describe a spiral staircase to someone who's never seen one. You can say the words, but they can't build the image in their head. The image token bypasses that problem entirely.

Jane: Exactly. And the paper actually has a great way of showing this. They gave the model the ground-truth image, the perfect imagination, and then asked it to answer. On Path Tracing, accuracy jumped to eighty-six point seven percent. That's a huge jump from the fifty percent when the model had to generate its own image.

Tom: So that tells us the bottleneck isn't the reasoning, it's the imagination quality. If the model can imagine the scene correctly, it can answer correctly. The problem is that generating that image is hard.

Meng: And that's where I want to ask a practical question, if I can. You've got this great result, but what does it cost? Generating images is computationally expensive. Is this something that can actually run in a real product, or is it just a research curiosity?

Tom: That's a fair question, Meng. And the paper has a good answer. They found that you don't need to generate the image at test time to get the benefit. The model trained with images does better even when it just answers directly. So you can train with the expensive method and then deploy a model that's fast at inference.

Meng: That's a relief. So the training cost is higher, but the deployment cost is the same as a standard model, just with better accuracy. That makes it practical.

Lu: And it also opens up a really interesting research direction. If we can get the imagination quality higher, the accuracy should go up too. The paper shows there's still a big gap between what the model imagines and the ground truth. Closing that gap is the next big challenge.

Jane: So we've got a method that works, we know why it works, and we know where the remaining challenges are. Let's wrap this up and think about what it all means.

Conclusion: Jane: Alright, Tom, let's bring it home. We've spent the show talking about "Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models," and I think we've only scratched the surface.

Tom: We really have. And I think the core message is simple. If you want AI to reason about space, you need to let it think in space. Text is a terrible medium for describing geometry. Images are the natural language of spatial reasoning.

Jane: And the paper shows this isn't just a nice idea. It's a measurable improvement. They beat the best closed-source models on Path Tracing. They show that training with imagination transfers to other tasks they weren't even trained on. This is a real step forward.

Lu: And I think the long-term impact is even bigger. This isn't just about counting chairs or navigating rooms. It's about building models that can construct mental models of the world. That's a foundational capability for robotics, for augmented reality, for any system that needs to understand and interact with physical space.

Meng: And from my side, the fact that you can get the benefit without the inference cost makes this something that could actually ship. It's not just a lab experiment. It's a practical technique.

Tom: So we're saying goodbye to this paper, but the ideas in it are going to stick with us. The next time you see a robot that can navigate a room it's never seen, or an AR system that can predict what's around a corner, you'll know where the seed was planted.

Jane: And that's the exciting part. This paper isn't the end of the story. It's the beginning of a new way to think about how AI understands space. Thanks for joining us, everyone. We'll see you next time with another paper to break down.

Tom: Take care, folks.

Mahtab Bigverdi, Linjie Li, Weikai Huang, Yiming Liu, Jaemin Cho, Tuhin Kundu, Chris Dongjoo Kim, Zelun Luo, Jieyu Zhang, Linda Shapiro, Ranjay Krishna

University of Washington · Allen Institute for AI · Microsoft · OpenAI

cs.AI

Submitted: 2026-08-17

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: Vision-language models (VLMs) excel at many tasks, yet continue to struggle with spatial reasoning—problems where the key information is not directly observable in the input.

Key concepts

Spatial Reasoning Difficulty
Current AI models excel at object recognition but fail when asked how a scene changes based on movement or perspective. They lack the ability to build a complete mental model of a space that goes beyond what is immediately visible in the input image.
Imaginative Perception Tokens
This is a training technique where the the model generates an internal image (a 'perception token') of what it believes it would see from a new viewpoint. This forces the AI to construct a mental scene before answering, allowing it to better understand complex spatial relationships.
Text Chain-of-Thought
This method requires the model to explain its reasoning using words before providing an answer. The hosts noted that for geometric tasks, this approach often hurts performance because describing a viewpoint rotation in words makes it difficult for models to build the necessary mental picture.

Terminology

Summary

Vision-language models (VLMs) excel at many tasks, yet continue to struggle with spatial reasoning—problems where the key information is not directly observable in the input. Many spatial questions require imaginative perception: simulating an unseen viewpoint, tracing a trajectory through an occluded space, or integrating partial views into a coherent spatial map. Humans naturally support this kind of reasoning through imagination. Prior work has introduced intermediate visual representations (e.g., visual thoughts, depth, or box tokens), but these intermediates often refine structure already visible rather than predicting the missing spatial structure implied by the evidence. We introduce Imaginative Perception Tokens (IPT), intermediate perceptual representations that externalize what a VLM would perceive under an alternative spatial configuration while remaining consistent with the observed input. To study this capability, we formulate three tasks that require imaginative perception: Perspective Taking (PET), Path Tracing (PT), and Multiview Counting (MVC). For each task, we construct datasets of ∼20K examples spanning simulated and real-world settings, paired with ground-truth intermediate imaginations, final answers, and curated evaluation benchmarks. Using the unified VLM BAGEL [12] as our backbone, IPT supervision improves spatial reasoning across several settings and often outperforms textual chain-of-thought training, even when no image is generated at inference time. For example, on MVC, IPT improves accuracy by 3.4% and achieves performance competitive with strong closed-source models on Path Tracing. We also find that mixed training with IPT and label-only data can further improve performance. In contrast, textual chain-of-thought can be detrimental on these tasks, substantially degrading performance in some cases, highlighting a modality mismatch when forcing spatial computation through language. Overall, IPT provides a principled supervision signal for reasoning over unobserved structure, yielding stronger spatial generalization and a more interpretable intermediate aligned with the underlying geometry of the task. Code will be released at the project page.

A key reason for this difficulty is that many spatial reasoning problems cannot be solved by analyzing the input alone. Instead, they require constructing a spatial representation that is not directly observed. Humans naturally address such problems through imagination: when asked what lies to the left after moving to a new position, or how many objects exist in a room seen from several viewpoints, we mentally simulate the scene from unseen perspectives or integrate partial observations into a unified spatial map [35, 43, 44]. In other words, spatial reasoning often depends on imagining missing spatial structure that proceed despite incomplete observations.

Existing approaches provide only partial solutions. Recent work teaches models to generate intermediate visual thoughts alongside language [15, 17, 21], while others introduce structured perceptual intermediates, such as depth maps or bounding boxes represented as tokens [2, 28, 40]. Although these methods demonstrate that intermediate visual representations can support reasoning, they primarily operate over information already present in the input observation, refining visible structures or extracting perceptual attributes. However, as discussed above, many spatial reasoning problems arise precisely because the required spatial information is not directly observable, and therefore requires imagination.

To address this gap, we propose Imaginative Perceptual Tokens for VLMs. When VLMs are trained with them, they enable intermediate reasoning steps that represent novel spatial views. Unlike standard perceptual intermediates that describe structures visible in the input, imaginative representations correspond to what the model would perceive if it were observing the input from a different spatial configuration, such as from an unseen viewpoint or after integrating multiple partial observations into one. At the same time, they are not unconstrained imagination: the predicted percept must remain consistent with the observed scene. These tokens externalize the model’s prediction of what would be perceived given incomplete spatial evidence.

To study this capability, we propose three spatial reasoning tasks that fundamentally require imaginative perception. (1) Perspective Taking requires predicting how a scene would appear from a new viewpoint given a single first-person observation (“If you move to the marked position and turn left, will the chair appear on your left or right?”); (2) Path Tracing requires inferring what an agent would see along a navigation path based on a top-down view (“If you walk along the marked path, which object will you see on your side?”); Finally, (3) Multiview Counting requires integrating multiple partial observations into a top-down view to determine the number of objects present in the scene. These tasks would be made easy when correctly predicting what would be perceived in a different spatial configuration. For each task we construct a dataset of approximately 20k examples each drawn from both real-world and synthetic simulated environments, with ground-truth intermediate spatial imaginations paired with final answers. Each dataset is accompanied by a human-filtered benchmark for evaluation. Together these constitute the first datasets designed explicitly to train and evaluate visually-grounded intermediate spatial reasoning in models.

Empirically, we find that training with imaginative perceptual supervision can improve performance on these spatial reasoning tasks compared to answer-only supervision, and often compares favorably to textual chain-of-thought approaches. These improvements can persist even when the model does not explicitly generate intermediate images at inference time, suggesting that such supervision may help models develop stronger internal spatial representations. At the same time, we observe that the benefits vary across tasks and settings, indicating that imagination quality and task structure both play important roles.

Overall, our results suggest that supervising models with intermediate perceptual predictions offers a useful direction for improving spatial reasoning, particularly in settings where the required structure is not directly observable from the input.

The core of our approach is to enable Multimodal Language Models (MLLMs) to externalize spatial reasoning through Imaginative Perception Tokens. Unlike standard textual chain-of-thought or methods outsourcing visual imagination with an external visual generation model, our method requires the model to generate a visual representation of a non-observed spatial configuration—such as an unseen viewpoint or an integrated top-down map—as a functional prerequisite for answering a spatial query.

Given an input context C consisting of one or more observed images Iobs = I1,..., Ik and a spatial language query Q, the goal is to predict the correct answer A. We decompose this into a two-stage generative process. First, the model generates imaginative perception tokens Iˆimag, representing the implied spatial structure requested by the task (e.g., the view from a new coordinate): P (Iˆimag Iobs, Q) Second, the conditioned on this imaginative perception tokens Iˆimag, the model produces the final answer: P (AIobs, Q, Iˆimag).

We implement this approach using BAGEL [12], a unified decoder-only transformer that natively supports interleaved multimodal understanding and generation. BAGEL employs a Mixture-of-Transformer-Experts (MoT) design: the model utilizes two transformer experts, one optimized for multimodal understanding and another for generation. Both operate on the same token sequence through shared self-attention at every layer. Images are represented via two distinct paths. Understanding tokens (U) are extracted via a SigLIP2 [31] ViT encoder to capture semantic content, while Generation tokens (G) are latent representations from a FLUX VAE used for high-fidelity synthesis. Because all tokens (text, U, and G) coexist in a single shared context window, the model maintains lossless interaction between understanding and generation modules. While BAGEL’s standard generation tokens are typically used for open-ended text-to-image generation or editing, we repurpose this generative capacity for spatial reasoning. In our framework, the generation target is not a stylistic output but a precise view imagination—a visually grounded intermediate that represents the unobserved 3D structure of the scene.

We optimize the framework using a multi-task loss Ltotal = λf m Lf m + λlm Llm. The model is trained to jointly produce the imaginative perception and the final answer: (1) Flow-Matching Loss (Lf m): For the imaginative intermediate, BAGEL adopts the Rectified Flow method. The model learns to predict the velocity field vt required to transform Gaussian noise into the target latent Ggt representing the unobserved view, conditioned on the preceding context C: Lf m = Et,G0,C ∥vt (Gt C) − (Ggt − G0)∥2. (2) Language Modeling Loss (Llm): We minimize the negative log-likelihood of the final VQA answer tokens A, conditioned on the observed context and the ground-truth imaginative tokens: Llm = − PA i=1 log P (ai C, Ugt, Ggt, a<i).

At inference time, the model operates in one of two modes depending on the task and configuration. In the text-only mode, the model produces only a textual answer without generating any visual intermediate A ∼ P (A C), serving as a baseline. In the imagination mode, the model first performs iterative denoising over VAE tokens to produce the imaginative latent: Ĝimag = R1 0 vt (Gt C) dt The decoded image Iˆimag is immediately re-encoded and appended to the context as both ViT understanding tokens and VAE generation tokens: C ′ = C, ViT(Iˆimag), VAE(Iˆimag) The model then attends to its own imagination to predict the final answer A ∼ P (A C ′).

We evaluate imaginative perception tokens on the three spatial reasoning tasks introduced in Sec. 3: Perspective Taking (PET), Path Tracing (PT), and Multiview Counting (MVC). To enable controlled comparisons, we train all task-specific models on the AI2-THOR subset of each dataset. We additionally report transfer to cross-environment benchmarks (Habitat), real-world images, and external datasets. All tasks use multiple-choice evaluation with balanced answer distributions.

We compare against two groups of models, evaluated zero-shot with task-specific prompts. VQA models include GPT-5, GPT-5.2, Gemini 2.5 Flash, Gemini 3 Flash, InternVL3.5-8B, Qwen2.5-VL-7B, and Qwen3-VL-8B. Unified models that support both understanding and generation include Janus-Pro-7B and Chameleon 7B. Our model variants include: Bagel (base): pretrained model with no task-specific fine-tuning; Bagel (label-only): fine-tuned with answer supervision only, with no intermediate thought; + Text CoT: trained to generate a textual chain-of-thought describing the imagined spatial configuration before answering; + IPT: trained to generate an intermediate image (the imaginative perception token) before answering; + Mixed Training: trained on a mixture of IPT examples (image-generation targets) and label-only examples (answer supervision only).

Spatial reasoning remains difficult for current VLM and unified models. Among the zero-shot baselines, GPT-5 is the strongest across nearly all settings, yet still trails our best fine-tuned variants on multiple in-distribution tasks. Smaller open VLM models (InternVL3.5-8B, Qwen2.5-VL-7B, Qwen3-VL-8B) hover near chance on PET (50–52%) and struggle on PT, indicating that these tasks are not solvable through superficial cues. Unified models perform worse overall: Chameleon 7B drops to 34.3% on PET and 5.4% on MVC, suggesting that current unified designs often trade away understanding robustness in exchange for generation capability.

Answer supervision alone yields large gains and transfers across environments. Bagel (label-only) substantially improves over Bagel (base) across all tasks, rising from 40.3% to 97.5% on AI2-THOR PET, from 29.9% to 65.7% on PT, and from 35.4% to 63.9% on MVC. These improvements transfer: label-only reaches 82.0% on Habitat PET, showing that spatial reasoning can be learned in simulation and generalized to new environments.

Imagination supervision helps most when language is a poor interface. On MVC, IPT achieves the best accuracy (67.3%), outperforming label-only (63.9%) and Text CoT (62.3%). On different-environment PET (Habitat), IPT reaches 87.0% (vs. 82.0% for label-only), and Mixed Training improves further to 87.7%. On PT, Mixed Training achieves the best results on both synthetic (66.7%) and real (58.6%) benchmarks, outperforming label-only (65.7% / 54.7%) and all baselines. IPT also improves real-world PT transfer (57.5%) over label-only (54.7%) and Text CoT (52.2%). Notably, IPT models are evaluated in answer-only mode: the model does not generate an image at inference, yet the imagination targets during training strengthen internal spatial representations that transfer across environments.

Text CoT underperforms label-only and IPT. Text CoT typically falls behind label-only (e.g., PET 83.1% vs. 97.5%, PT 49.7% vs. 65.7%) and also behind IPT (e.g., MVC 62.3% vs. 67.3%, PET 83.1% vs. 96.8%). Compared to label-only, the Text CoT objective forces the model to allocate capacity to generating long spatial descriptions during fine-tuning, which competes with answer prediction. Compared to IPT, the gap reflects a modality mismatch: viewpoint changes, occlusions, and cross-view correspondences are difficult to serialize into natural language, and the resulting textual traces introduce noise rather than useful structure. IPT represents these relationships directly in the visual modality where they are naturally expressed.

Latent resolution controls imagination quality and downstream accuracy. At Latent-4 (64 × 64), imaginations are blurry and lose spatial detail; at Latent-64 (1024 × 1024), imaginations become sharper and more spatially faithful, preserving object identities and relative positions. Quantitatively, increasing resolution from Latent-4 to Latent-64 improves AI2-THOR PET from 87.4% to 96.8% and MVC from 53.5% to 63.1%. Habitat PET peaks at Latent-32 (87.0%) and drops slightly at Latent-64 (83.3%), suggesting mild overfitting to AI2-THOR appearance statistics at the highest resolution.

IPT training builds stronger spatial representations than Text CoT. On PT, IPT with answer-only inference (61.1%) outperforms Text CoT with answer-only inference (55.8%) by 5.3 points. On MVC, IPT with image generation (63.1%) outperforms Text CoT with text generation (61.5%). Imagination supervision is useful, but explicit generation is not required at inference. For IPT models, answer-only mostly outperforms generating the imagination explicitly: on PT, answer-only reaches 61.1% vs. 50.4% with generation. For Text CoT, generating the chain-of-thought also slightly underperforms answer-only (53.1% vs. 55.8% on PT), though the gap is smaller than for IPT. This asymmetry suggests that producing faithful imaginations is harder than producing text descriptions, and imperfect generations can mislead downstream reasoning. However, training with imagination targets remains valuable: answer-only IPT matches GPT-5 on PT (61.1%). Ground-truth imaginations reveal headroom. When given ground-truth imaginations instead of model-generated ones, PT accuracy jumps from 50.4% to 86.7% (+36.3) and MVC rises from 63.1% to 67.3% (+4.2). The large PT gap indicates that imagination quality is the dominant bottleneck for path tracing; for PET, model-generated imaginations nearly match GT (96.8% vs. 96.7%), leaving little room for improvement.

IPT transfers to aligned external benchmarks. On SAT, Bagel (label-only) improves from 34.9% to 59.1% over Bagel (base), and Mixed Training further improves to 63.6%. On MessyTable, Mixed Training reaches 37.0%, up from 29.0% for Bagel (base). Training with our data improves performance on other spatial benchmarks. Bagel (fine-tuned) consistently improves over Bagel (base) across all three benchmarks (40.5%→52.0% on ScanNet, 39.5%→47.5% on MindCube, 40.0%→50.0% on All-Angles), indicating that our simulator data builds broadly useful spatial representations even when the specific imaginative token target changes.

We introduced Imaginative Perception Tokens (IPTs), intermediate visual representations that externalize spatial reasoning about unobserved structure in multimodal language models, and designed three tasks: Perspective Taking, Path Tracing, and Multiview Counting, with ground-truth intermediate imaginations. Training with imagination supervision consistently outperforms label-only and text chain-of-thought baselines, even without explicit imagination at inference, and ablations confirm that imagination quality directly governs downstream accuracy.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement and the resulting capabilities of the improved AI system:

  • Implementation: Extend the training loss to include a flow-matching objective (L fm) that supervises the model to generate a latent representation of an unobserved spatial view (e.g., novel viewpoint, sideview, or top-down BEV map) before answering.

  • Specifics: Use a unified decoder-only architecture (like BAGEL) with Mixture-of-Transformer-Experts. During training, condition generation on the input context and optimize L total = λ fm * L fm + λ lm * L lm, where L fm predicts the velocity field for denoising Gaussian noise into the target latent, and L lm is the standard answer token loss.

  • Implementation: Replace text-based CoT supervision with IPT supervision for tasks involving viewpoint changes, occlusions, or cross-view integration. The model learns to generate a visual intermediate (e.g., a rotated view) rather than a textual description.

  • Specifics: For tasks like Is the chair on my left after turning 90°?, train the model to first generate the imagined view from the new position, then answer. This avoids the modality mismatch where spatial geometry is poorly serialized into language.

  • Implementation: Increase the resolution of the imaginative perception token from Latent-4 (64×64) to Latent-64 (1024×1024) during training.

  • Specifics: This improves downstream accuracy by up to 9.4% (e.g., from 87.4% to 96.8% on Perspective Taking) because higher-resolution imaginations preserve object identities and relative positions needed for reasoning.

  • Implementation: Train on a 50/50 mixture of examples with IPT supervision (visual generation targets) and examples with only answer supervision. Use different system prompts to switch modes.

  • Specifics: This yields the best performance across tasks (e.g., 97.8% on PET, 66.7% on PT, 62.3% on MVC) and improves transfer to real-world benchmarks (e.g., 63.6% on SAT vs. 59.1% for label-only).

  • Implementation: After training with IPT supervision, at inference time, allow the model to skip explicit image generation and directly output the answer.

  • Specifics: This matches or exceeds the performance of generating the image first (e.g., 61.1% vs. 50.4% on Path Tracing), showing that imagination training strengthens internal spatial representations without requiring externalization.

  • Capability: Given a first-person image with a marked target position, the system can correctly determine if an object becomes closer/further or appears on the left/right after moving and turning, even though the target view is never shown.

  • Performance: Achieves 96.8% accuracy on AI2-THOR and 87.0% on Habitat (unseen environments), outperforming GPT-5 (79.8%) and all other baselines.

  • Capability: Given a top-down map with a marked path and egocentric endpoint views, the system can infer which object is visible on a queried side at a midpoint, requiring 3D visibility reasoning not present in the input.

  • Performance: Achieves 66.7% on synthetic and 58.6% on real-world Matterport3D benchmarks, competitive with or better than GPT-5 (60.2% synthetic).

  • Capability: Given 4+ first-person frames of a scene, the system can generate a top-down bird's-eye view map and count unique objects, resolving occlusions and cross-view duplicates.

  • Performance: Achieves 67.3% on AI2-THOR, outperforming label-only (63.9%) and text CoT (62.3%), and transfers to real-world MessyTable (37.0% with mixed training).

  • Capability: The system generalizes from synthetic training (AI2-THOR) to photorealistic (Habitat), real photos (Matterport3D, MessyTable), and abstract geometric benchmarks (MindCube, All-Angles-Bench).

  • Performance: Improves from 40.5% to 52.0% on ScanNet, 39.5% to 47.5% on MindCube, and 40.0% to 50.0% on All-Angles, showing broad spatial capability gains.

  • Capability: The system can externalize its spatial reasoning as a viewable image (novel viewpoint, sideview, or BEV map), making its reasoning process inspectable and verifiable by humans.

  • Performance: Generated imaginations at Latent-64 are sharp and spatially faithful, enabling users to see what the model thinks the scene looks like from an unobserved perspective.

  • Capability: Unlike text CoT, which can degrade performance by up to 14.4% (e.g., PET 83.1% vs. 97.5% label-only), the IPT-trained system maintains high accuracy because spatial relationships are computed in the visual modality where they are naturally expressed.

Abstract

Vision language models (VLMs) excel at many tasks but still struggle with spatial reasoning when critical information is not directly observable. Many such problems require imaginative perception: inferring what would be seen from an unseen viewpoint, tracing paths through occluded spaces, or integrating partial observations into a coherent spatial representation. We introduce Imaginative Perception Tokens (IPT), intermediate perceptual representations that externalize what a VLM would perceive under alternative spatial configurations while remaining consistent with the observed input. To study this capability, we formulate three tasks, Perspective Taking (PET), Path Tracing (PT), and Multiview Counting (MVC), and construct datasets of approximately 20K examples with ground truth imaginations, answers, and evaluation benchmarks. Using the unified VLM BAGEL as the backbone, IPT supervision consistently improves spatial reasoning and often outperforms textual chain of thought training, even without generating images at inference time. On MVC, IPT improves accuracy by 3.4% and achieves competitive performance with strong closed-source models on PT. We further find that combining IPT and label-only supervision yields additional gains, whereas textual chain of thought can substantially degrade performance, suggesting a modality mismatch when spatial computation is forced through language. Overall, IPT provides a principled supervision signal for reasoning about unobserved spatial structure, improving generalization while producing interpretable intermediate representations.

Sources

Related papers