LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

summary

Video file (mp4)

The gist

LEGO-Puzzles is a benchmark designed to systematically evaluate the spatial reasoning capabilities of Multimodal Large Language Models (MLLMs), from basic spatial understanding to multi-step planning.

In short

The episode discusses 'LEGO-Puzzles,' a paper testing MLLMs' multi-step spatial reasoning using LEGO bricks. Hosts analyze that while models perform moderately on basic tasks, they fail drastically when required to plan or generate intermediate steps, concluding current AI lacks sequential spatial understanding.

Key concepts

MLLMs
Multimodal large language models (MLLMs) are types of AI capable of looking at an image and answering questions about it. The paper uses these models to test their ability to understand and reason about three-dimensional space.
Multi-Step Spatial Reasoning
This is the ability for an AI model to not only see objects but also plan a sequence of actions or steps in time and space, such as figuring out how to build a structure piece by piece.
Elementary Set
The first task set in the paper, consisting of over 1,100 questions across eleven tasks. These range from simple questions (e.g., which object is taller) to identifying the next piece in an assembly.
Planning Set
The second and more complex task set where models are given a starting LEGO structure and a final target. They must then correctly identify and order the necessary intermediate steps, scaling up to eight steps.

Terminology used across episodes

This episode discusses

The paper

LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning? · Read on arXiv

Yanhong Zeng, Haodong Duan, Junyao Gao, Kexian Tang, Yanan Sun, Zhening Xing, Wenran Liu, Kai Chen, Kaifeng Lyu

Tsinghua University · Tongji University · Shanghai AI Laboratory

Many real-world applications of spatial intelligence, such as robotic control, autonomous driving, and automated assembly, require spatial reasoning across multiple sequential steps. However, the extent to which current Multimodal Large Language Models (MLLMs) possess this capability remains largely unexplored. Inspired by LEGO construction, a recreational activity that critically relies on multi-step spatial reasoning, we introduce LEGO-Puzzles: a benchmark designed to systematically evaluate the spatial reasoning capabilities of MLLMs from basic spatial understanding to multi-step planning. LEGO-Puzzles contains two task sets. The Elementary set covers 11 visual question-answering (VQA) tasks with 1,100 carefully curated samples to test elementary spatial reasoning skills that are cruical for LEGO assembly. The Planning set directly requires the model to generate a step-by-step plan for assembling a target LEGO structure, where the tasks are organized into subsets with different planning horizons ranging up to 8. Our evaluation of 29 state-of-the-art MLLMs shows that even the strongest models struggle with elementary reasoning tasks in LEGO construction, falling at least 20% behind human performance. The planning accuracy also quickly drops to 0% as the number of planning steps increases, whereas our human participants solve all the tasks perfectly. Switching the output format from multiple choice to image generation degrades model performance even further, leading to zero accuracy even for planning 3 steps. Overall, LEGO-Puzzles reveals critical limitations in current MLLMs' spatial reasoning capabilities and highlights the need for substantial advances.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?".

Jane: The paper was written by Yanhong Zeng, Haodong Duan, Junyao Gao, Kexian Tang, Yanan Sun et al. from Tsinghua University and Tongji University and Shanghai AI Laboratory.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone! Today we’re cracking open a paper that’s got a fun title and a serious question behind it: “LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?” Jane, I have to say, when I first saw the title, I thought, okay, we’re gonna watch AI build little plastic castles. But it’s so much more than that.

Jane: Absolutely, Tom. And for anyone just tuning in, MLLMs are multimodal large language models—the kind of AI that can look at an image and answer questions about it. The authors took something we all played with as kids, LEGO bricks, and turned it into a rigorous test. They wanted to know: can these models actually understand three dee space, and can they reason through a sequence of steps to build something?

Tom: Right, and the title is almost cheeky because it sounds simple, but the results are pretty humbling. I mean, we’re talking about models like GPT-five and Gemini-two point five-Pro, the big names, and they’re getting beaten by humans by over twenty percent on the basic tasks.

Jane: That’s the kicker, Tom. The paper is from a team at Tsinghua University, Tongji, and the Shanghai AI Lab, and they didn’t just make a toy benchmark. They built a systematic evaluation. The title says “How Good Are MLLMs,” and the answer so far is: not as good as we hoped, especially when you ask them to plan several steps ahead.

Tom: And that’s what we’re going to dig into today. The implications here stretch way beyond toys—think robots assembling things, or AI helping you follow IKEA instructions. If they can’t handle LEGO, we’ve got a long way to go.

Jane: Exactly. So stick around, because we’re going to break down the two big task sets, the numbers, and what this means for the future of spatial AI. Tom, I think the next segment is where we really get into the meat of the summary.

Tom: Let’s do it. We’ve got a lot to unpack.

Summary: Tom: So, Jane, let’s get into the actual summary of “LEGO-Puzzles.” The paper sets up two main test sets. The first is the Elementary set—over one thousand one hundred questions across eleven tasks. These range from simple stuff like “which object is taller in three dee space?” to harder stuff like “which piece comes next in this assembly?”

Jane: And the second set is the Planning set, which is where it gets really interesting. They give the model a starting LEGO structure and a final target, and the model has to pick the correct intermediate steps and put them in order. They scale it from one step all the way up to eight steps.

Tom: Right, and the results are brutal. On the Elementary set, the best model, GPT-five scored seventy-two percent. Humans scored ninety-three point six percent on a smaller subset. That’s a massive gap. But the Planning set is where it gets almost comical. At eight steps, GPT-five drops to zero percent accuracy. Zero. And humans get one hundred percent across the board.

Jane: It’s not just about picking the right pieces either. The paper also tests whether models can generate images of the intermediate steps. So instead of multiple choice, the model has to actually draw the next state of the LEGO build. And that’s even worse. At three steps, the accuracy hits zero for the best models.

Tom: Yeah, that’s the part that really got me. These models can generate beautiful images of a cat wearing a hat, but they can’t generate a slightly taller LEGO tower. The spatial precision just isn’t there.

Jane: And that’s the core finding, Tom. The paper isn’t just saying “AI is bad at LEGO.” It’s showing that current MLLMs lack a fundamental capability: sequential spatial reasoning. They can recognize objects, they can name colors, but they can’t hold a mental model of a three dee scene and update it step by step.

Tom: So when we talk about implications, this is huge for robotics, for augmented reality, for any field where a machine has to manipulate physical objects. But we’ll get into that more. For now, I want to talk about what the paper suggests we should do about it.

Jane: Good, because the paper doesn’t just point out the problem—it offers a path forward.

Improvements: Tom: Alright, so the paper suggests some clear improvements, and I think this is where the conversation gets constructive. Jane, what did you take away from their recommendations?

Jane: Well, the first thing they emphasize is that we need benchmarks that actually scale in complexity. A lot of existing tests only check single-step reasoning—like, “is this object to the left of that one?” But LEGO-Puzzles forces models to chain multiple steps together. The authors argue that we need more of this progressive difficulty to really push the field forward.

Tom: And they also point out something about evaluation methods. They tried using GPT-4o as a judge to score the generated images, and it didn’t match human judgment well. So they’re saying: we can’t just rely on AI to grade AI for these spatial tasks. We need human evaluation, or at least better automated metrics.

Jane: That’s a really important point. If we can’t reliably measure progress, we can’t make progress. They also highlight that switching from multiple-choice to open-ended generation is a much harder test. And that’s a good thing—it exposes weaknesses that multiple-choice questions hide.

Tom: Right, because in multiple choice, the model can sometimes guess or use subtle cues. But when it has to generate the next state from scratch, there’s nowhere to hide. The paper suggests that future models need to be trained not just to recognize spatial relationships, but to imagine and render them.

Jane: And that’s a big ask. It means we need new training objectives, maybe more three dee data, maybe more interactive environments. The authors don’t prescribe a specific method, but they make it clear that current approaches are hitting a wall.

Tom: So the improvements are really about two things: better benchmarks to expose the gaps, and better evaluation to measure them honestly. But I think the deeper question is what this means for real-world applications. Let’s bring in Lu and Meng for that.

Jane: Good idea. Let’s get their take on the practical side.

First Page Discussion: Tom: So we’re looking at the first page of “LEGO-Puzzles” again, and I want to pull out something specific. The authors mention that LEGO construction is used in cognitive science as a reliable indicator of spatial intelligence. That’s a nice hook, but it also grounds the whole paper in something real.

Jane: It does, Tom. And on that first page, they lay out the core problem: we have lots of benchmarks for spatial understanding, but very few for sequential spatial reasoning. That’s the gap they’re filling. They’re not just testing whether a model can see—they’re testing whether it can think in time and space.

Lu: If I can jump in here, Jane, that’s exactly what excites me. The first page frames this as a stepping stone toward real-world applications like robotic control and automated assembly. When a robot picks up a block, it doesn’t just need to know where the block is—it needs to know where to put it next, and then what to do after that. This benchmark is a way to test that chain of decisions.

Meng: And from my side, the engineering challenge is huge. The paper shows that even the best models fail at generating correct intermediate states. That means if you’re building a system that uses an MLLM to guide a user through, say, fixing a piece of furniture, it’s going to give you wrong instructions after the second step. That’s a reliability problem we can’t ignore.

Tom: That’s a great point, Meng. And it connects back to the first page’s emphasis on “multi-step planning.” The authors are saying that understanding a single image isn’t enough. You need to understand a sequence of images and how they transform.

Jane: Right, and the first page also introduces the two task sets, Elementary and Planning, which we’ve already covered. But what I find striking is the confidence with which they state the results—that even the strongest models fall at least twenty percent behind humans. That’s a bold claim, and the data backs it up.

Lu: It does, and I think that’s the most valuable contribution here. They’ve created a benchmark that is hard enough to expose real limitations, but structured enough that we can see exactly where the failures happen. That’s going to guide a lot of future research.

Tom: So we’ve got the problem, the benchmark, and the results. Let’s wrap this up with what it all means.

Conclusion: Tom: Alright, let’s bring it home. We’ve spent the show talking about “LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?” and honestly, the title is the perfect summary. The answer is: not good enough yet.

Jane: That’s right. The paper gives us a clear picture. On basic spatial understanding, models like GPT-five and Gemini-two point five-Pro do okay, but they’re still far behind humans. And when you ask them to plan multiple steps, they fall apart completely. At eight steps, the best model scores zero.

Lu: And the generative part is even more telling. When models have to create the intermediate images themselves, they fail at just three steps. That tells us that current MLLMs are great at recognizing patterns but terrible at maintaining a consistent three dee mental model over time.

Meng: For anyone building real systems, that’s a warning sign. You can’t trust these models for tasks that require precise, sequential physical reasoning—not yet. But that’s also an opportunity. This benchmark gives us a clear target to improve against.

Tom: And that’s the silver lining. The paper doesn’t just say “AI is bad.” It provides a tool to measure progress. Future models can be tested against LEGO-Puzzles to see if they’re actually getting better at spatial reasoning.

Jane: So we say goodbye to this paper with a sense of excitement, not disappointment. The gap is real, but now we can see it clearly. And seeing the problem is the first step to solving it.

Tom: Well said, Jane. Thanks to Lu and Meng for joining us, and to our listeners for sticking around. Next time, we’ll pick up another paper and see what other mysteries of AI we can unpack. Until then, keep building.

Jane: And keep questioning. See you all next time.

More episodes

← Home