TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning

arXiv:2601.16520 · cs.CV, cs.AI, cs.CL · Submitted 2026-01-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning".

Jane: The paper was written by N/A (Authors not found in provided excerpts) from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

The Summary: Tom: So, what's the core of this paper? They introduce TangramPuzzle as a geometry-grounded benchmark designed to expose these gaps in compositional spatial reasoning.

Jane: They use the classic Tangram game, which uses seven fixed pieces that we have to assemble into a target shape using only translations and rotations.

Meng: But here's the critical innovation: they call it the Tangram Construction Expression or TCE, which represents every single instance with exact, machine-verifiable coordinate specifications.

Lu: That TCE is huge because it completely eliminates any ambiguity that comes from relying on fuzzy visual approximations; we’re talking about mathematical rigor here.

Lalam: It moves the evaluation from an artistic judgment to a verifiable calculation of how well the AI can handle concrete spatial rules.

Tom: And this rigorous framework powers two distinct tasks, which is what makes the test so comprehensive.

Jane: The first is Outline Prediction, where we show the final shape and ask models to infer what they are made of based on local components.

Meng: Then the End-to-End Code Generation task requires solving the inverse problem—taking a target silhouette and figuring out exactly how to decompose it into seven pieces.

Lu: It's like asking the AI to not only look at a finished painting but also reverse engineer how every single brushstroke was applied.

Jane: It’s a powerful way to test that compositional skill, moving from simple pattern matching to true spatial construction.

The Improvements and Findings: Tom: We've talked about the structure, but the findings are where it gets interesting—the results show a systemic failure mode in current MLLMs.

Jane: They found that models tend to prioritize matching the final target silhouette, which is visually appealing, but they neglect the actual geometric constraints of how those pieces should fit together.

Lu: That's a real cognitive bias; the visual goal overrides the mathematical requirement for maintaining rigid body integrity.

Meng: It’s a huge practical finding because it means that if we want an AI to perform complex tasks like automated assembly or surgical planning, simply making it look at pictures isn't enough.

Lalam: The tendency to "cheat" by distorting pieces is worrying because it suggests the model is optimizing for visual fidelity rather than structural correctness.

Tom: But while their performance was generally low, the results were quite clear about which models succeeded and which struggled.

Jane: Gemini3-Pro really stands out here, achieving near-ceiling performance with an accuracy of ninety-eight point six five percent on the Outline Prediction task alone.

Lu: It seems like that level of success is tied to its ability to handle complex reasoning without visual shortcuts, which is a testament to its architecture.

Meng: The fact that it's so much better suggests that for complex spatial tasks, we need models with a truly robust understanding of geometric logic, not just high-level pattern recognition.

Jane: So, the paper is showing us exactly how these modern AI models fail when they are being asked to be precise about their physical actions.

Implications: Tom: This brings up massive implications for the future development of AI in fields that require physical action.

Jane: If we want robots or autonomous systems to perform complex tasks, like assembling a machine or even just navigating a cluttered space, they need this compositional spatial reasoning ability.

Meng: If the AI can't reliably solve this puzzle, it cannot be used for precise manufacturing or sophisticated assembly tasks where the parts must fit perfectly.

Lu: It suggests that we need to move past benchmarks that are purely semantic and require a simple "yes" or "no" answer, and start testing our models against these complex, solvable physical problems.

Lalam: On a cultural level, it shows us how far we are from truly reliable machine intelligence; we’ still rely on visual shortcuts when the stakes require mathematical precision.

Tom: It's not just about robots either failing to assemble a toy; it's about any task where the AI must decompose a complex scene into its constituent parts and verify their fit.

Jane: It forces us to ask, "Is this model just seeing me, or is it truly understanding how the pieces relate to one another?"

Meng: If we are deploying these systems in real-world environments with physical objects, we cannot afford an AI that prioritizes a visual match over adhering to rigid geometric constraints.

Lu: We need models that can handle the exact coordinates and the algebraic expressions, not just ones that look plausible.

Conclusion: Tom: We've seen how TangramPuzzle is designed to be rigorously machine-verifiable, moving away from fuzzy visual approximations entirely.

Jane: It’s a great tool for evaluating whether AI can handle complex spatial reasoning or if it's just relying on superficial cues.

Lu: The evidence is clear that this compositional space is one of the most challenging areas for current models, and we need to keep pushing the boundaries of this kind of testbed.

Meng: I think this research provides a very concrete framework for how engineers can start building validation pipelines for future AI systems that require physical dexterity.

Lalam: It gives us a new way to measure not just intelligence, but the ability to achieve structural integrity and reliable execution in the real world.

Tom: And before we go, I want to hear one last thought from each of you.

Jane: This paper has successfully raised the bar for what it means to be "multimodal" in a truly meaningful way.

Meng: It makes my job easier because it gives us a clear, mathematical specification of failure modes that we can actively test against our models.

Lalam: I feel this pushes the AI towards greater responsibility by demanding accuracy over aesthetic appeal.

Lu: It forces the ultimate rigor on the computational models, ensuring they understand the physical constraints of reality itself.

Tom: Thank you all for sharing your insights into TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning.

Jane: We'll see you next time!

cs.CV, cs.AI, cs.CL

Submitted: 2026-01-23

Updated: 2026-09-07

Importance score: 82/100

The gist: This paper introduces the TangramPuzzle benchmark, a rigorous evaluation platform designed to assess "Multimodal Large Language Models with Compositional Spatial Reasoning." It establishes a

Key concepts

Tangram Puzzle
The puzzle uses seven fixed pieces that must be assembled into a target shape. The only allowed operations are translations and rotations of these pieces, making it a test of precise geometric manipulation.
Tangram Construction Expression (TCE)
TCE is a method that represents every puzzle instance using exact, machine-verifiable coordinate specifications. This mathematical rigor eliminates any ambiguity found in fuzzy visual approximations, ensuring the evaluation is based on precise calculation rather than artistic judgment.
Compositional Spatial Reasoning
This concept tests an AI's ability to go beyond simple pattern matching. It requires the model to understand how individual components relate spatially and can be assembled into a whole, moving from visual recognition to true spatial construction.
Outline Prediction
This task requires the AI to look at the final assembled shape (the silhouette) and infer what local components were used to create it. It tests the model's ability to reverse-engineer the structure of a complex scene.

Terminology

Summary

This paper introduces the TangramPuzzle benchmark, a rigorous evaluation platform designed to assess Multimodal Large Language Models with Compositional Spatial Reasoning. It establishes a demanding testbed for evaluating how well LLMs can perform complex, multi-step geometric assembly tasks by fitting constituent pieces to form a target silhouette. The benchmark is critical because it moves beyond simple visual recognition, requiring models to demonstrate physical feasibility and precise geometric adherence—skills that are essential but often poorly measured in current multimodal AI systems.

Annotation and Data Rigor

To ensure the high quality and rigorous geometric precision of the benchmark, the data annotation process involved inviting researchers with backgrounds in geometry and computer science. The annotation interface utilized a magnetic snapping mechanism that automatically avoid[ed] pixel-level misalignment. Furthermore, evaluation metrics account for various failure modes beyond simple matching. These include:

  • PE (Physical Error): Evaluates physical feasibility, flagging invalid predictions if pieces overlap or if the union of all pieces is not a single connected shape.

  • TSE (Syntax Error): Checks structural compliance with the required TCE format, flagging issues like unparsable JSON or incorrect piece counts.

  • Invalid Rate: Defined as the proportion of instances in which the model output does not yield any valid option that can be mapped to the candidate set.

Shape Similarity and Geometric Metrics

The assessment of shape similarity employs a combination of complementary metrics to capture both global overlap and local boundary fidelity. The primary metrics are:

  • IoU (Intersection over Union): Measures the overlap between the predicted assembly U and the target silhouette T, calculated as mu(U T) over mu(U T).

  • RGE (Rigid Geometry Error): Verifies that each predicted piece maintains its original shape by comparing its area and perimeter against the ground-truth specification.

  • Hausdorff Distance (dH): This metric captures boundary-level deviations by calculating the maximum distance between the boundaries of U and T, ensuring that dH(d U, d T) is minimized.

Impact of Context and Modality

The study investigates how providing textual context affects model performance, revealing a complex dependency on modality. When examining In-context Learning (ICL), the authors noted that while ICL yields clear benefits for silhouette-level quality in valid responses (improved IoU and reduced Hausdorff distances), it simultaneously caused models to exhibit increased Syntax Error (TSE) rates, suggesting that the symbolic context imposes a cognitive load that interferes with the model’s ability to strictly adhere to structural constraints.

Furthermore, an ablation study removing explicit textual descriptions of the target outline revealed a Stability Collapse. The elimination of text specifications led to a dramatic surge in syntax errors for models like GPT-5.2. This confirmed a widespread inability to precisely 'read' geometric coordinates directly from raw visual inputs, highlighting that structured generation is heavily dependent on semantic prompts, with Gemini3-Pro noted as an exception demonstrating advanced capability to perform rigorous reasoning solely based on visual grounding.

Improvements for AI systems

System Architecture Enhancement for Geometric Reasoning and Constraint Adherence

The existing models exhibit critical failure modes when transitioning from visual pattern recognition to mathematically rigorous physical assembly. The core improvement must involve decoupling the high-level semantic understanding (what the puzzle is) from the low-level, non-negotiable geometric verification (how it must fit).

Here are three specific, multi-layered improvements:


Improvement: Integrate a dedicated, differentiable geometric constraint solver module that operates after the initial VLM hypothesis generation but before final output structuring. This module must treat physical and topological constraints as hard loss functions rather than soft penalties.

Technical Specificity:

The GCSM will utilize a combination of Graph Neural Networks (GNNs) to model piece connectivity (nodes = pieces, edges = adjacency/contact points) and an Optimization Layer (e.g., using differentiable physics simulation) to enforce geometric feasibility.

  • Constraint Enforcement: The module must explicitly calculate and minimize the following losses:

  • L Overlap: A repulsive force potential applied between any pair of predicted piece polygons (P i, P j) if their intersection area mu(P i P j) > epsilon, forcing them apart until mu(P i P j) = 0.

  • L Connectivity: A penalty proportional to the number of disconnected components in the union of predicted pieces (PE enforcement).

  • L Boundary Adherence: A distance metric minimizing the Hausdorff distance between the assembled shape's boundary and the target silhouette d T during refinement steps.

Improved AI Capability:

The system gains Guaranteed Physical Validity. It moves beyond merely suggesting a shape to proving a physically realizable assembly. If an initial VLM output violates the non-overlap or single-connectedness constraints, the GCSM will iteratively adjust piece vertices using gradient descent against L Overlap and L Connectivity until a mathematically valid configuration is achieved or failure is proven. This directly addresses the Partial Visual Similarity through Impermissible Physical Violations observed in model predictions.

  • Process Flow: After the VLM generates raw tokens for the output structure, a preliminary parser attempts to map these tokens into an abstract syntax tree (AST) based on the required JSON schema.

  • Self-Correction: If parsing fails (e.g., incorrect key name, missing array bracket), the AST failure is fed back as a specific error signal (TSE error) directly to the VLM's attention mechanism via Prefix Tuning. The prompt context is temporarily augmented with: ERROR: JSON structure failed at key 'pieces'. Re-evaluate based on schema constraint.

  • Input Processing: The raw image is processed by a Vision Transformer (ViT) backbone that outputs localized feature maps. A secondary Convolutional Neural Network (CNN) branch is trained specifically to predict key geometric primitives:

  1. Vertex Coordinates: Predicting (x, y) pairs for every corner of the target silhouette d T.

  2. Edge Vectors: Predicting the vectors defining the edges connecting these vertices.

  • Attention Fusion: The feature vectors derived from this CNN/ViT combination are then used to condition and modulate the cross-attention layers of the main VLM encoder, effectively forcing the model to ground its reasoning in predicted coordinates rather than just semantic concepts.

Sources

Related papers