LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?".
Jane: The paper was written by Yanhong Zeng, Haodong Duan, Junyao Gao, Kexian Tang, Yanan Sun et al. from Tsinghua University and Tongji University and Shanghai AI Laboratory.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone! Today we’re cracking open a paper that’s got a fun title and a serious question behind it: “LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?” Jane, I have to say, when I first saw the title, I thought, okay, we’re gonna watch AI build little plastic castles. But it’s so much more than that.
Jane: Absolutely, Tom. And for anyone just tuning in, MLLMs are multimodal large language models—the kind of AI that can look at an image and answer questions about it. The authors took something we all played with as kids, LEGO bricks, and turned it into a rigorous test. They wanted to know: can these models actually understand three dee space, and can they reason through a sequence of steps to build something?
Tom: Right, and the title is almost cheeky because it sounds simple, but the results are pretty humbling. I mean, we’re talking about models like GPT-five and Gemini-two point five-Pro, the big names, and they’re getting beaten by humans by over twenty percent on the basic tasks.
Jane: That’s the kicker, Tom. The paper is from a team at Tsinghua University, Tongji, and the Shanghai AI Lab, and they didn’t just make a toy benchmark. They built a systematic evaluation. The title says “How Good Are MLLMs,” and the answer so far is: not as good as we hoped, especially when you ask them to plan several steps ahead.
Tom: And that’s what we’re going to dig into today. The implications here stretch way beyond toys—think robots assembling things, or AI helping you follow IKEA instructions. If they can’t handle LEGO, we’ve got a long way to go.
Jane: Exactly. So stick around, because we’re going to break down the two big task sets, the numbers, and what this means for the future of spatial AI. Tom, I think the next segment is where we really get into the meat of the summary.
Tom: Let’s do it. We’ve got a lot to unpack.
Summary: Tom: So, Jane, let’s get into the actual summary of “LEGO-Puzzles.” The paper sets up two main test sets. The first is the Elementary set—over one thousand one hundred questions across eleven tasks. These range from simple stuff like “which object is taller in three dee space?” to harder stuff like “which piece comes next in this assembly?”
Jane: And the second set is the Planning set, which is where it gets really interesting. They give the model a starting LEGO structure and a final target, and the model has to pick the correct intermediate steps and put them in order. They scale it from one step all the way up to eight steps.
Tom: Right, and the results are brutal. On the Elementary set, the best model, GPT-five scored seventy-two percent. Humans scored ninety-three point six percent on a smaller subset. That’s a massive gap. But the Planning set is where it gets almost comical. At eight steps, GPT-five drops to zero percent accuracy. Zero. And humans get one hundred percent across the board.
Jane: It’s not just about picking the right pieces either. The paper also tests whether models can generate images of the intermediate steps. So instead of multiple choice, the model has to actually draw the next state of the LEGO build. And that’s even worse. At three steps, the accuracy hits zero for the best models.
Tom: Yeah, that’s the part that really got me. These models can generate beautiful images of a cat wearing a hat, but they can’t generate a slightly taller LEGO tower. The spatial precision just isn’t there.
Jane: And that’s the core finding, Tom. The paper isn’t just saying “AI is bad at LEGO.” It’s showing that current MLLMs lack a fundamental capability: sequential spatial reasoning. They can recognize objects, they can name colors, but they can’t hold a mental model of a three dee scene and update it step by step.
Tom: So when we talk about implications, this is huge for robotics, for augmented reality, for any field where a machine has to manipulate physical objects. But we’ll get into that more. For now, I want to talk about what the paper suggests we should do about it.
Jane: Good, because the paper doesn’t just point out the problem—it offers a path forward.
Improvements: Tom: Alright, so the paper suggests some clear improvements, and I think this is where the conversation gets constructive. Jane, what did you take away from their recommendations?
Jane: Well, the first thing they emphasize is that we need benchmarks that actually scale in complexity. A lot of existing tests only check single-step reasoning—like, “is this object to the left of that one?” But LEGO-Puzzles forces models to chain multiple steps together. The authors argue that we need more of this progressive difficulty to really push the field forward.
Tom: And they also point out something about evaluation methods. They tried using GPT-4o as a judge to score the generated images, and it didn’t match human judgment well. So they’re saying: we can’t just rely on AI to grade AI for these spatial tasks. We need human evaluation, or at least better automated metrics.
Jane: That’s a really important point. If we can’t reliably measure progress, we can’t make progress. They also highlight that switching from multiple-choice to open-ended generation is a much harder test. And that’s a good thing—it exposes weaknesses that multiple-choice questions hide.
Tom: Right, because in multiple choice, the model can sometimes guess or use subtle cues. But when it has to generate the next state from scratch, there’s nowhere to hide. The paper suggests that future models need to be trained not just to recognize spatial relationships, but to imagine and render them.
Jane: And that’s a big ask. It means we need new training objectives, maybe more three dee data, maybe more interactive environments. The authors don’t prescribe a specific method, but they make it clear that current approaches are hitting a wall.
Tom: So the improvements are really about two things: better benchmarks to expose the gaps, and better evaluation to measure them honestly. But I think the deeper question is what this means for real-world applications. Let’s bring in Lu and Meng for that.
Jane: Good idea. Let’s get their take on the practical side.
First Page Discussion: Tom: So we’re looking at the first page of “LEGO-Puzzles” again, and I want to pull out something specific. The authors mention that LEGO construction is used in cognitive science as a reliable indicator of spatial intelligence. That’s a nice hook, but it also grounds the whole paper in something real.
Jane: It does, Tom. And on that first page, they lay out the core problem: we have lots of benchmarks for spatial understanding, but very few for sequential spatial reasoning. That’s the gap they’re filling. They’re not just testing whether a model can see—they’re testing whether it can think in time and space.
Lu: If I can jump in here, Jane, that’s exactly what excites me. The first page frames this as a stepping stone toward real-world applications like robotic control and automated assembly. When a robot picks up a block, it doesn’t just need to know where the block is—it needs to know where to put it next, and then what to do after that. This benchmark is a way to test that chain of decisions.
Meng: And from my side, the engineering challenge is huge. The paper shows that even the best models fail at generating correct intermediate states. That means if you’re building a system that uses an MLLM to guide a user through, say, fixing a piece of furniture, it’s going to give you wrong instructions after the second step. That’s a reliability problem we can’t ignore.
Tom: That’s a great point, Meng. And it connects back to the first page’s emphasis on “multi-step planning.” The authors are saying that understanding a single image isn’t enough. You need to understand a sequence of images and how they transform.
Jane: Right, and the first page also introduces the two task sets, Elementary and Planning, which we’ve already covered. But what I find striking is the confidence with which they state the results—that even the strongest models fall at least twenty percent behind humans. That’s a bold claim, and the data backs it up.
Lu: It does, and I think that’s the most valuable contribution here. They’ve created a benchmark that is hard enough to expose real limitations, but structured enough that we can see exactly where the failures happen. That’s going to guide a lot of future research.
Tom: So we’ve got the problem, the benchmark, and the results. Let’s wrap this up with what it all means.
Conclusion: Tom: Alright, let’s bring it home. We’ve spent the show talking about “LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?” and honestly, the title is the perfect summary. The answer is: not good enough yet.
Jane: That’s right. The paper gives us a clear picture. On basic spatial understanding, models like GPT-five and Gemini-two point five-Pro do okay, but they’re still far behind humans. And when you ask them to plan multiple steps, they fall apart completely. At eight steps, the best model scores zero.
Lu: And the generative part is even more telling. When models have to create the intermediate images themselves, they fail at just three steps. That tells us that current MLLMs are great at recognizing patterns but terrible at maintaining a consistent three dee mental model over time.
Meng: For anyone building real systems, that’s a warning sign. You can’t trust these models for tasks that require precise, sequential physical reasoning—not yet. But that’s also an opportunity. This benchmark gives us a clear target to improve against.
Tom: And that’s the silver lining. The paper doesn’t just say “AI is bad.” It provides a tool to measure progress. Future models can be tested against LEGO-Puzzles to see if they’re actually getting better at spatial reasoning.
Jane: So we say goodbye to this paper with a sense of excitement, not disappointment. The gap is real, but now we can see it clearly. And seeing the problem is the first step to solving it.
Tom: Well said, Jane. Thanks to Lu and Meng for joining us, and to our listeners for sticking around. Next time, we’ll pick up another paper and see what other mysteries of AI we can unpack. Until then, keep building.
Jane: And keep questioning. See you all next time.
Yanhong Zeng, Haodong Duan, Junyao Gao, Kexian Tang, Yanan Sun, Zhening Xing, Wenran Liu, Kai Chen, Kaifeng Lyu
Tsinghua University · Tongji University · Shanghai AI Laboratory
cs.AI
Submitted: 2026-08-09
Updated: 2026-08-11
Code: https://github.com/opendatalab/PDF-Extract-Kit
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 69/100
The gist: LEGO-Puzzles is a benchmark designed to systematically evaluate the spatial reasoning capabilities of Multimodal Large Language Models (MLLMs), from basic spatial understanding to multi-step planning.
Key concepts
- MLLMs
- Multimodal large language models (MLLMs) are types of AI capable of looking at an image and answering questions about it. The paper uses these models to test their ability to understand and reason about three-dimensional space.
- Multi-Step Spatial Reasoning
- This is the ability for an AI model to not only see objects but also plan a sequence of actions or steps in time and space, such as figuring out how to build a structure piece by piece.
- Elementary Set
- The first task set in the paper, consisting of over 1,100 questions across eleven tasks. These range from simple questions (e.g., which object is taller) to identifying the next piece in an assembly.
- Planning Set
- The second and more complex task set where models are given a starting LEGO structure and a final target. They must then correctly identify and order the necessary intermediate steps, scaling up to eight steps.
Terminology
Summary
LEGO-Puzzles is a benchmark designed to systematically evaluate the spatial reasoning capabilities of Multimodal Large Language Models (MLLMs), from basic spatial understanding to multi-step planning. Inspired by LEGO construction, the benchmark contains two task sets. The Elementary set covers 11 visual question-answering (VQA) tasks with 1,100 carefully curated samples to test elementary spatial reasoning skills that are crucial for LEGO assembly. The Planning set directly requires the model to generate a step-by-step plan for assembling a target LEGO structure, where the tasks are organized into subsets with different planning horizons ranging up to 8.
The Elementary set is divided into three levels. Level 1 (Spatial Understanding) includes tasks on Height, Adjacency, Rotation, and Multiview. Level 2 (Single-Step Sequential Reasoning) includes Rotation Status, Position, Next-Step, and Dependency. Level 3 (Multi-Step Sequential Reasoning) includes Backwards, Ordering, and Outlier. The Planning set includes two tasks: Plan-k-Step-VQA and Plan-k-Step-Generation, which evaluate the model's ability to generate a multi-step plan in VQA and image generation settings, respectively.
The evaluation of 29 state-of-the-art MLLMs shows that even the strongest models struggle with elementary reasoning tasks in LEGO construction, falling at least 20% behind human performance. The planning accuracy also quickly drops to 0% as the number of planning steps increases, whereas human participants solve all the tasks perfectly. Switching the output format from multiple choice to image generation degrades model performance even further, leading to zero accuracy even for planning 3 steps.
Key findings include: a clear gap between proprietary and open-source models, with GPT-5 achieving 72.0% overall accuracy while most open-source models perform only marginally better than random guessing. Humans outperform the best MLLMs by over 20%, with human experts achieving 93.6% overall performance on a subset of the benchmark. MLLMs struggle with elementary tasks such as Height and Rotation, where many models perform below random guessing. In the Planning set, all models exhibit a clear decline in exact match accuracy as the number of planning steps increases, with GPT-5 dropping from 90% at k=1 to 0% at k=8. Even when partial correctness is considered, model performance exceeds random guessing by less than 20%.
For image generation tasks, even the strongest current MLLMs fail to perform full-stage generative planning, with errors often appearing from the very first step. Human evaluation of single-turn image generation tasks shows that proprietary models outperform open-source ones in both appearance consistency and instruction adherence, but all models show clear room for improvement in instruction following. The paper also discusses the correlation between LEGO-Puzzles and real-world spatial reasoning, finding strong positive correlations with 3DSRBench tasks (0.93 for Height and 0.98 for Adjacency), validating its utility as a proxy for evaluating real-world spatial understanding.
The paper concludes that LEGO-Puzzles reveals critical limitations in current MLLMs' spatial reasoning capabilities and highlights the need for substantial advances. The main contributions are a comprehensive evaluation framework with progressively increasing spatial and sequential complexity, natural but challenging planning tasks constructed from open-source human-designed LEGO projects, and evaluation of multi-turn image generation where all frontier MLLMs fail completely even for 3 planning steps.
Improvements for AI systems
Based on the paper, I can implement the following specific improvements to AI systems:
Improvement: Add explicit 3D-aware training data and loss functions that penalize 2D-projection-based answers.
What the improved system can do:
-
Correctly judge relative heights of objects in 3D space, even when 2D projections are misleading (e.g., objects at different depths appearing similar heights)
-
Distinguish between objects that are adjacent vs. separated in 3D, not just visually touching in 2D
-
Accurately determine rotation angles (30°, 60°, 90°, 120°) of objects around their centers
Abstract
Many real-world applications of spatial intelligence, such as robotic control, autonomous driving, and automated assembly, require spatial reasoning across multiple sequential steps. However, the extent to which current Multimodal Large Language Models (MLLMs) possess this capability remains largely unexplored. Inspired by LEGO construction, a recreational activity that critically relies on multi-step spatial reasoning, we introduce LEGO-Puzzles: a benchmark designed to systematically evaluate the spatial reasoning capabilities of MLLMs from basic spatial understanding to multi-step planning. LEGO-Puzzles contains two task sets. The Elementary set covers 11 visual question-answering (VQA) tasks with 1,100 carefully curated samples to test elementary spatial reasoning skills that are cruical for LEGO assembly. The Planning set directly requires the model to generate a step-by-step plan for assembling a target LEGO structure, where the tasks are organized into subsets with different planning horizons ranging up to 8. Our evaluation of 29 state-of-the-art MLLMs shows that even the strongest models struggle with elementary reasoning tasks in LEGO construction, falling at least 20% behind human performance. The planning accuracy also quickly drops to 0% as the number of planning steps increases, whereas our human participants solve all the tasks perfectly. Switching the output format from multiple choice to image generation degrades model performance even further, leading to zero accuracy even for planning 3 steps. Overall, LEGO-Puzzles reveals critical limitations in current MLLMs' spatial reasoning capabilities and highlights the need for substantial advances.
Sources
- GPT-4 Technical Report
- Pixtral 12B
- Qwen3-VL Technical Report
- Building and better understanding vision-language models: insights and future directions
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
- Qwen2.5-VL Technical Report
- X-IQE: eXplainable Image Quality Evaluation for Text-to-Image Generation with Visual Large Language Models
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Gemini: A Family of Highly Capable Multimodal Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Emu3: Next-Token Prediction is All You Need
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- OpenAI GPT-5 System Card
- GamiBench: Evaluating Spatial Reasoning and 2D-to-3D Planning Capabilities of MLLMs with Origami Folding Tasks
- Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection