Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

arXiv:2608.30751 · cs.AI, cs.CV · Submitted 2026-08-31 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models".

Jane: The paper was written by Ashwin Nedungadi, Stefan Oehmcke and Stefan Ludtke from Institute for Visual and Analytical Computing, University of Rostock.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: The core of the paper lies in introducing Autoregressive Mosaics or AM-Bench, which separates two major factors that previous research always mixed together: spatial composition and spatial expression. They needed a way to isolate whether a model *know* how to arrange things versus just knowing how to *write the code*to make them look right.

Jane: It’s important that we don’re not just looking at a final image because, as they point out, an LLM-generated picture alone can't tell us if it failed because it couldn't arrange the pieces or if it just couldn't write the specific programming language to achieve that arrangement.

Lu: The summary of the results is really telling. They found that across eight different open-weight models, while all reliably translated a fully specified geometry into code, their ability to compose an image from an underspecified prompt—the layout task—was vastly different. That's where the real difference in spatial reasoning emerged.

Meng: The fact that the translation task acts as a control is crucial for me. It ensures that when we see a model fail on the layout task, we know it’s not because it’s fundamentally incapable of coding, but because something else is going wrong with its spatial planning.

Lalam: This suggests that AI isn't just mimicking human visual ability; it's engaging in a process of 'compositional thought.' That compositional power could revolutionize how we generate complex scenes for storytelling and visual media.

Tom: It’s clear the authors have established this separation, which sets us up perfectly to discuss the improvements they suggest.

Improvements: Tom: Moving past the initial findings, let's look at what the paper suggests as key insights or "improvements" to our understanding of these models. One major finding is that code generation ability isn't the main bottleneck for spatial performance; layout composition itself is a harder challenge.

Jane: And it’s not just about the model either medium matters, which they demonstrate by replacing procedural canvas code with raw SVG generation and seeing performance improve across all models. This shows us the output medium is an active constraint on what we can achieve.

Lu: The most insightful finding, in my view, is that a coarse spatial layout plan is actually present before the model generates any code at all, but it only reflects the prompt's implied layout. That’s a huge step toward understanding internal representation.

Meng: My practical concern here is that if you could improve models by using raw SVG instead of their native canvas API, that would be a major engineering win for development cycles in visual AI applications.

Lalam: The implication of the model tracking the evolving geometric state during generation, rather than executing a fixed plan, suggests that we can build more adaptive and dynamic AI systems that could create truly emergent art.

Tom: We've covered all our key findings on how these models operate and where they might be improved, which is a great transition into our final wrap-up.

Conclusion: Tom: As we wrap up this segment, let’s summarize the overall conclusion of Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models. The paper concludes that the performance of text-only AI models on spatial tasks depends on two distinct factors.

Jane: First, their ability to express a usable internal layout exists, and second, their ability to express that layout through the output medium they are given. Neither factor alone explains all the observed performance differences across various AI models.

Lu: I think it's significant that this work shows a generalized spatial reasoning capability in these massive models, even if only a coarse plan is present at the start of any generation process. It points to emergent abilities that we need to study further.

Meng: The practical insight here is that when building next-generation AI tools, we shouldn't just assume code fluency solves all problems; we need to consider the limitations imposed by the rendering medium and how our prompts structure spatial constraints.

Lalam: Ultimately, this suggests a pathway toward achieving embodied spatial intelligence in AI, allowing us to build systems capable of creating complex visual narratives that truly reflect internal understanding.

Tom: It's been a fantastic discussion about Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models. Thank you all for sharing your insights today.

Lu: I’m excited to see how this opens up the possibilities for next time around will be.

Meng: We're already looking into ways to use these findings to make our systems more efficient and robust.

Lalam: And I hope we can share this kind of emergent intelligence with the public in a way that enhances our shared culture.

Conclusion: Tom: So we’ve spent a lot of time breaking down "Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models," and the general idea is that AI models are capable of spatial reasoning, but their abilities aren't monolithic.

Jane: That’s a huge point, Tom. We learned that simply knowing how to write code doesn' not explain why some models are better at composing a picture than others.

Lu: I agree with Jane; the distinction is crucial because it shows we aren't just seeing random pattern matching in these LLMs.

Meng: From an engineering viewpoint, it’s interesting that this suggests two different paths for improvement: better planning and better expression through the medium.

Lalam: The insight that they are tracking an evolving geometric state rather than a fixed plan is what I find most impactful for our future work on how AI can represent complex visual information.

Tom: It's definitely not just about code generation, as the paper’s translation task proves, but it's also not purely about spatial composition, which is what the layout task adds.

Jane: The authors really succeeded in separating those two factors to give us a much clearer picture of what we are observing in these models.

Lu: And it seems like this distinction could be tested across various other emergent abilities too, not just 2D mosaics.

Meng: I'm looking forward to seeing how we can apply this concept of designing for the medium rather than just seeing the model as a monolithic black box.

Lalam: The ability to move beyond a static plan into an incremental, autoregressive process is a step towards giving AI more dynamic creative potential.

Tom: That’s a great way to put it, Lalam. We've seen that "Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models" gives us a lot of ground to stand on.

Jane: It’s clear this is just the beginning of many questions about these models, and I think it’s exciting to hear what other research will reveal.

Lu: Absolutely; we are only scratching the surface of how these architectures actually work internally.

Meng: We're ready to see how these principles translate into real-world production systems.

Lalam: And I hope that this understanding leads to a new era of richer and more spatially aware digital content for our culture.

Ashwin Nedungadi, Stefan Oehmcke, Stefan Ludtke

Institute for Visual and Analytical Computing, University of Rostock

cs.AI, cs.CV

Submitted: 2026-08-31

Updated: 2026-09-01

Comments: Pre-Print

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 86/100

The gist: The paper introduces AM-Bench, a novel benchmark designed to measure "allocentric 2D layout at mosaic granularity." This framework serves as a "necessary-condition probe for embodied spatial

Key concepts

Autoregressive Mosaics (AM-Bench)
A framework introduced in the paper designed to separate two factors in AI: spatial composition (knowing how to arrange things) and spatial expression (writing the code to make them look right). This allows researchers to isolate a model's true understanding of space.
Layout Task
The task where models must compose an image from an underspecified prompt. The hosts emphasize that performance on this task is crucial because it reveals the difference in a model's spatial planning ability, separate from its coding skills.
Spatial Composition vs. Spatial Expression
The core distinction made by the paper. Spatial composition refers to a model's inherent ability to understand and arrange objects in space, while spatial expression is the model's ability to write the specific code required for that arrangement.

Terminology

Summary

The paper introduces AM-Bench, a novel benchmark designed to measure allocentric 2D layout at mosaic granularity. This framework serves as a necessary-condition probe for embodied spatial competence by evaluating how well text-only language models can perform complex visual reasoning tasks. The study is significant because it systematically tests whether modern generative AI models possess inherent spatial planning capabilities, moving beyond simple linguistic understanding to assess structured, geometric comprehension.

Scope and Benchmark Design

AM-Bench measures 2D layout by requiring the model to generate autoregressive mosaics, which are structural compositions of shapes. The benchmark is designed to be comprehensive, covering various levels of spatial difficulty. The test suite includes multiple types of challenges:

  1. Iconic Prompts: These prompts require recognizing and reproducing complex visual structures, such as A traffic light (tall dark rectangle, th...) or A yellow sun with eight radiating rays o...

  2. Periodicity/Tiling: The benchmark specifically tests for periodicity—repeating grids or tilings—an axis that was historically difficult to build into language models.

  3. Compositional Inversion: This task assesses the model's ability to handle complex structural relationships, such as an Iconic/Compositional inversion.

Evaluation Metrics and Analysis

The evaluation utilizes multiple metrics to quantify spatial performance, including the J1 score (a measure of layout quality) and attention standard deviation. The analysis also employs token-generation entropy to gauge consistency. For instance, Figure 16 illustrates the relationship between entropy and local-structure distance across eight different models, showing a strong correlation (Spearman rho = 0.93). Performance is measured against various model architectures and sizes, including Qwen, Gemma, GLM, and Llama/CodeLlama families (8B–34B).

Observed Model Behaviors and Limitations

The study reports distinct behavioral patterns across the tested models. For example, in the DINO attention maps (Figure 17), the top model (GLM-4 32B) often demonstrates highly focused attention on key structural elements, whereas the worst model (CodeLlama 34B) shows less precise localization. However, researchers caution that these findings are not definitive: it is still not a controlled scaling-law sweep, and whether the observed patterns hold at larger scales or for closed-weight models remains untested.

Future Directions and Technical Depth

To further validate its claims, the research proposes several avenues for expansion. These include:

  • Scaling the benchmark to encompass broader spatial categories through a larger multiannotator human validation run.

  • Introducing other output mediums as an additional code-prior ablation to investigate the effect of the output medium.

  • Applying advanced causal analysis techniques, specifically activation patching for a direct, causal analysis of the models’ internal spatial planning, to substantiate behavioral claims at the representation level.

The authors also provide extensive technical detail regarding data quality, noting that some Layout samples are missing a second-judge score and that the pooled foreground mask can be empty for approximately 3% of structure-coherence cells when describing several small, spatially diffuse elements.

Improvements for AI systems

The provided text excerpts detail sophisticated, multi-faceted evaluations of current LLMs, focusing heavily on spatial reasoning (Layout/Iconic prompts), code generation reliability (Entropy analysis), and multimodal understanding (DINO attention). The gaps highlighted—especially regarding periodicity, metric continuity, and the limited scope of the current benchmarks—represent critical areas for immediate architectural and methodological improvement.

Here are the specific, high-impact improvements required to advance AI systems using this research foundation.


Improvement 1: Implementing Causal Activation Patching for Internal Spatial Planning.

  • Technical Action: Do not rely solely on surface-level output metrics (like Jaccard Index or layout scores). Integrate and utilize activation patching techniques during inference to perform a direct, causal analysis of the model's internal spatial representations. Specifically, target the transformer layers responsible for compositionality and object permanence within the visual embedding space.

  • Rationale: The paper notes that behavioral claims should be substantiated at the representation level. We must move beyond correlation to causation when assessing spatial reasoning failures (e.g., why does it fail on periodicity?).

  • Improved System Capability: The AI system can generate an Explainable Spatial Failure Report. When a layout task fails, the system doesn't just output Wrong. It identifies which specific internal attention head or activation subspace failed to maintain object coherence (e.g., Failure point: Loss of relative depth metric between object A and background B in Layer 12). This allows for targeted architectural retraining.

Improvement 2: Integrating a Metric-Aware, Multiscale Representation Module.

  • Technical Action: Modify the backbone architecture to explicitly model metric egocentric continuity, depth, and occlusion as first-class citizens of the latent space. This requires coupling the standard ViT encoder with a dedicated depth estimation module (e.g., using a monocular depth prediction network) that is jointly trained with the layout task loss function (L total = L layout + lambda 1 L depth + lambda 2 L metric continuity).

  • Rationale: The current benchmark scope is too narrow (does not cover metric egocentric continuity, depth, or occlusion). A robust AI must understand 3D spatial relationships from 2D prompts.

  • Improved System Capability: The AI can perform True Embodied Spatial Planning. Given a prompt like "Place the chair behind the table and slightly above the potted plant," it generates not just a plausible image, but an internal, consistent 3D coordinate map that respects relative depth and occlusion rules, making it suitable for robotic control or advanced simulation environments.

Improvement 3: Developing and Mandating the Periodicity/Tiling Stress-Test Ladder.

  • Technical Action: Formalize the novel periodicity rung into a mandatory, scaled component of all spatial evaluation suites. This set must include tasks that require understanding repeating patterns, rotational symmetries, and tiling constraints (e.g., tessellations with specific edge matching rules).

  • Rationale: The paper explicitly notes that periodicity is the most likely axis of degradation and remains untested. Failure here indicates a fundamental weakness in pattern recognition beyond simple compositionality.

  • Improved System Capability: The AI can demonstrate Advanced Pattern Synthesis and Prediction. It moves beyond merely reproducing a tiled image to generating the missing component or predicting the next state in a repeating, constrained pattern, proving deep understanding of geometric rules rather than rote memorization.

Improvement 4: Implementing a Comprehensive Noise-Aware Target Set Scoring System.

  • Technical Action: Formalize the handling of ambiguous target inputs (like the 17 targets needing simplification). The evaluation pipeline must incorporate an automated pre-processing step that scores the ambiguity of the target description (e.g., using graph theory metrics on touching/overlapping shapes) and penalizes models based on their failure to generalize across varying levels of noise/ambiguity.

  • Rationale: Source noise is a documented weakness (0.96 vs 1.00). A production-grade system cannot fail simply because the prompt was slightly vague or contained minor overlaps that required human simplification.

  • Improved System Capability: The AI achieves Robust Zero-Shot Interpretation. It can process and generate accurate output even when the input prompts are imperfect, noisy, or require implicit merging of adjacent elements (merging touching, same-color shapes), making it reliable in real-world data streams.

Improvement 5: Integrating Code Generation Entropy as a Pre-Filter and Fine-Tuning Loss.

  • Technical Action: The entropy measurement must transition from being merely an analysis tool to a direct training objective. Implement a loss function component (L entropy) during fine-tuning that penalizes the model when its generated code deviates significantly from the expected low-entropy, high-consistency pattern observed in valid samples.

  • Rationale: The entropy trend (Figure 16) shows a strong correlation between local structure distance and generation consistency. Using this as a loss term forces the model to learn predictable internal representations for structured data (like code or geometric coordinates).

  • Improved System Capability: The AI exhibits Guaranteed Code and Structure Consistency. When asked to generate complex, multi-step code or coordinate data, it not only produces syntactically correct output but also guarantees that the underlying structure is maximally consistent and predictable, minimizing runtime bugs caused by internal drift.

Improvement 6: Developing a Matched-Scale Distillation Baseline for Low-Resource Scenarios.

  • Technical Action: Implement a structured distillation pipeline where the full, large-scale model (the teacher, e.g., Qwen2.5-Coder 32B) is tasked with generating high-quality, verified examples across difficult domains (Iconic/Compositional inversions). These samples are then used to train a smaller, highly efficient student model on a matched scale, ensuring the student retains the full behavioral capability of the teacher without requiring its vast parameter count.

  • Rationale: This addresses the necessity of deploying high-capability models in resource-constrained environments while maintaining state-of-the-art performance.

  • Improved System Capability: The AI achieves High Performance, Low Latency Deployment. It can execute complex spatial reasoning tasks and generate sophisticated code with minimal computational overhead, making it viable for edge devices or high-throughput API services where latency is critical.

Abstract

Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whether this reflects an internal representation of 2D spatial layout or simply the ability to translate spatial descriptions into code. We introduce Autoregressive Mosaics (AM-Bench), a benchmark that separates these factors: First, a translation task gives a model a fully specified geometry of a picture in words as a prompt and asks for the code that produces it. Second, a layout task requires the model to compose an image from an underspecified prompt. Across eight open-weight text-and-code-only models, all models reliably translate specified geometry into code, but their open-ended layout performance differs substantially, indicating that these differences are not explained by code-generation ability alone. An output-medium ablation further shows that the interface or medium of expression that the model uses matters: replacing procedural code with raw SVG improves layout scores across all models. Finally, probing model activations shows that a coarse layout plan is present before generation, but reflects only the layout implied by the prompt. During generation, models track the evolving geometric state instead of executing an initially fixed plan. Overall, these results show that 2D spatial performance in text-only LLMs depends on both the model and the output medium, and is not explained by code-generation ability alone.

Sources

Related papers