Can AI Understand the Language of Origami?

arXiv:2603.13856 · cs.LG, cs.CV · Submitted 2026-03-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Can AI Understand the Language of Origami?".

Jane: Building AI systems capable of planning and acting in physical environments requires understanding causal mechanisms governing physical processes, which necessitates internal representations that link observations, actions, and environmental changes.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's look at the specific title and the people behind this work. "Can AI Understand the Language of Origami?" it’s a very direct question about whether an AI can learn to speak the language of physical transformation through folding instructions.

Jane: The authors are Yihan Wu, Yichang Jian, and Yifei Peng Yao-Xiang Ding from the State Key Lab of CAD andCG at Zhejiang University, and Nishad Mansoor from Northeastern University. They bring a strong background in both computer science and geometric modeling to this problem.

Lu: Their expertise seems perfectly suited because they are working on areas that require understanding structured physical transformations, which is exactly what origami embodies one <ref:2603.13856#pg0>.

Meng: I’m curious how their specific research background helps them tackle the challenge of moving from visual input to a precise sequence of folds and then verifying if those folds actually work in the real world.

Lalam: The team’s focus seems to be on creating an interactive benchmark, OrigamiBench, which is a way to test these models by forcing them into a closed-loop system where they propose actions and get immediate feedback on physical validity.

The paper's summary: Tom: Now that we know the setup, let's talk about what the paper actually summarizes. They are proposing OrigamiBench, which is an interactive environment designed to test if AI models can reason about geometric transformations while synthesizing shapes through folding operations.

Jane: Essentially, they set up a scenario where an agent has to take a flat sheet of paper and fold it into a target 2D shape using folds that must be geometrically valid <ref:2603.13856#pg0>. The key is that the process involves recursive composition, which is very similar to how programming works for building complex structures eight <ref:2603.13856#pg0>.

Lu: The summary highlights that the folding process naturally captures structured physical transformations grounded in visual perception, making origami a compact domain for testing this kind of reasoning one <ref:2603.13856#pg0>.

Meng: They detail the environment as a closed-loop system where the agent proposes folds and gets feedback on whether those folds are physically valid or how similar they are to the target configuration. That iterative feedback loop is crucial for training.

Lalam: The paper summarizes that experiments with modern vision–language models showed three key findings: first, scaling model size alone doesn't reliably produce causal reasoning about transformations <ref:2603.13856#pg1>, second, models struggle to generate coherent multi-step folding plans because their visual and language representations aren't fully integrated <ref:2603.13856#pg1>, and third, there's a big gap between performance when models just rely on visual similarity versus when they have to infer the underlying folding operations, which is what we call causal reasoning <ref:2603.13856#pg1>.

The paper's improvements: Tom: So, if that summary is accurate, what are the actual suggested improvements the authors propose for these AI systems? They aren't just saying "it doesn't work well," they’re suggesting concrete ways to fix it.

Jane: The authors suggest strengthening the connection between language, perception, and geometry by introducing explicit intermediate state representations. They talk about things like crease graphs or vertex–edge structures that models can use as a guide one <ref:2603.13856#pg0>.

Lu: I think the suggestion to explicitly represent these geometric structures is vital because it forces the model to operate with more structured information rather than just processing raw pixels, which should help bridge that gap in causal reasoning.

Meng: From an engineering perspective, this sounds like a way to give the AI a better internal map of the physical state, which would make planning much more reliable for tasks requiring sequential decision-making World Action Planner.

Lalam: Another improvement they suggest is leveraging the closed-loop simulator as a learning environment for execution-based supervision through reinforcement or interactive learning with rewards based on foldability and geometric validity. That shifts the training from just getting the final answer to actually succeeding in the physical process.

Conclusion: Tom: So, wrapping up this discussion on "Can AI Understand the Language of Origami?", we see that while current vision–language models show some ability, scaling them up isn't automatically making them good at causal reasoning about physical transformations <ref:2603.13856#pg1>.

Jane: The paper emphasizes that the real direction for progress involves building systems where the language understanding is explicitly grounded in the visual perception of geometric transformations one <ref:2603.13856#pg0>. It’s about getting those visual and symbolic concepts to talk to each other better.

Lu: I think this focus on intermediate representations, like crease graphs, is where we need to concentrate our efforts if we want AI to start reasoning about physical constraints more deeply one <ref:2603.13856#pg0>.

Meng: For practical applications, this means future AI systems could be much more reliable when tasked with generating physically constructible designs because they'd be checking validity at every step rather than just at the end Trade-off Functions for DP-SGD with Subsampling based on Random Allocation.

Lalam: I think the implication here is that by focusing on these causal mechanisms, we move beyond simple pattern matching and toward systems that can actually plan and act in the physical world successfully <ref:2603.13856#pg0>.

Tom: Well said, Lalam. It sounds like this work gives us a clear roadmap for how to push AI toward deeper causal understanding in physical tasks. We'll be looking forward to seeing how these new representations are put into practice in the next set of research papers we cover.

Naaisha Agarwal, Yihan Wu, Yichang Jian, Yifei Peng Yao-Xiang Ding, Nishad Mansoor, Yikuan Hu Mohan Li Wang-Zhou Dai, Emanuele Sansone

State Key Lab of CAD&CG, Zhejiang University · Computer Science Department, Northeastern University · National Key Laboratory for Novel Software Technology, Nanjing University · CSAIL / ESAT / MIT / KU Leuven

cs.LG, cs.CV

Submitted: 2026-03-14

Updated: 2026-10-02

Comments: This version: "Can AI Understand the Language of Origami?" - different paper from v1 with different authors - NeurIPS LP4FM (Outstanding Runner-Up Award) v1: OrigamiBench: An Interactive Environment to Synthesize Flat-Foldable Origamis ICML LM4Plan (Oral)

Code: https://github.com/origamimagiro/flatfolder

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: Building AI systems capable of planning and acting in physical environments requires understanding causal mechanisms governing physical processes, which necessitates internal representations that

Key concepts

OrigamiBench
An interactive benchmark designed to test if AI can reason about geometric transformations while synthesizing 2D shapes from a flat sheet of paper via folding. It involves iteratively proposing folds and receiving feedback on physical validity and similarity to a target configuration.
CreasePattern object
The internal state of the environment, stored as a structured .fold file. This object precisely tracks all geometric details including vertex coordinates, edge connectivity, fold types (Mountain or Valley), and face definitions. It allows the system to maintain an accurate representation of the paper's folded state.
Causal vs. Associative Setting
The difference between two ways models are tested. The associative setting relies on visual similarity (pattern matching). The causal setting requires models to infer underlying physical folding operations and understand cause-and-effect to achieve a target shape, which is where current AI struggles.

Terminology

Summary

Building AI systems capable of planning and acting in physical environments requires understanding causal mechanisms governing physical processes, which necessitates internal representations that link observations, actions, and environmental changes. This capability is tested by evaluating models on benchmarks that jointly assess visual perception and structured reasoning about physical transformations in domains like origami.

OrigamiBench Introduction

The paper introduces OrigamiBench as an interactive benchmark designed to evaluate whether AI models can reason about geometric transformations while synthesizing shapes through folding operations. The core challenge lies in constructing shapes from a flat sheet of paper into a target 2D shape via a sequence of folds constrained by geometric validity and foldability. This domain is compact enough for reasoning about geometric transformations, yet open-ended enough to support the evaluation of generalization and creative problem solving, reflecting recursive composition key to programming.

OrigamiBench Environment

The environment operates as a closed-loop system where the agent iteratively proposes folds and receives feedback on physical validity and similarity to a target configuration. The workflow follows a strict sequence: the environment presents a multimodal input, the agent responds with a structured command, and the environment validates and executes it to update its state. The internal state is maintained as a CreasePattern object, which can be serialized into a.fold file format capturing vertex coordinates, edge connectivity, fold type assignments (Mountain or Valley), and face definitions.

Data Organization

The origami dataset comprises 366 distinct designs encoded in the standardized.fold file format. Each sample is organized along two dimensions: semantic category and complexity level. The semantic category captures the representational type of the model, with classes including roles, animals, plants, geometry, and complexity level (Easy, Medium, or Hard) is determined by structural complexity based primarily on vertex and crease line counts.

Evaluation Tasks & Metrics

OrigamiBench is organized around two primary tasks: a one-step evaluation and a full-step interactive evaluation. The first task assesses perception, understanding the geometric transformation of a single fold, and predicting the outcome. It is formulated as a multiple-choice classification problem where the model must identify which candidate state can be obtained by performing EXACTLY ONE forward fold operation. The second task, full-step interactive evaluation, assesses planning capabilities by requiring models to synthesize target origami shapes starting from a blank crease state through sequential actions. Performance in this setting is evaluated using three metrics: (i) Query Efficiency (QE), defined as the percentage of fold steps that contribute to the final origami; (ii) Geometric Similarity (GS), computed using Intersection over Union to quantify geometric overlap between the model-generated and target shapes; and (iii) Semantic Similarity (SS), measured as cosine similarity between vector embeddings obtained from a fine-tuned CLIP model.

Experimental Findings

Experiments with modern vision–language models revealed three key findings. First, scaling model size alone does not reliably produce causal reasoning about transformations. Second, models struggle to generate coherent multi-step folding plans, suggesting visual and language representations remain weakly integrated. Third, the results indicate a significant gap between performance on the Associative setting (where models rely on visual similarity) and the Causal setting (where models must infer underlying folding operations), indicating current VLMs struggle to move beyond pattern matching toward genuine causal reasoning about geometric transformations.

Future Work

The paper suggests that strengthening the coupling between language, perception, and geometry is a promising direction by introducing explicit intermediate state representations (e.g., crease graphs, vertex–edge structures) and training models to jointly predict states and actions. Another direction involves leveraging the closed-loop simulator as a learning environment for execution-based supervision through reinforcement or interactive learning with rewards derived from foldability and geometric validity. The authors also plan to extend diagnostic analyses through challenge splits involving more complex crease patterns and tighter geometric constraints to better promote causal understanding.

The gist: Scaling model size alone does not reliably produce causal reasoning about transformations, models struggle to generate coherent multi-step folding plans, and visual and language representations remain weakly integrated.

How it works

  1. The environment maintains a precise internal state implemented as a CreasePattern object, serialized to a.fold file format detailing vertex coordinates, edge connectivity, fold type assignments (Mountain or Valley), and face definitions.

  2. The agent receives structured observations delivered as multimodal prompts including visual inputs (target shape, current folded state, crease pattern) and binary feedback on the feasibility of the previous action.

  3. The agent outputs a single structured action defined as a JSON object specifying an action like add crease, detailing edge vertices and assignment ("M or V").

  4. Upon receiving an action, the environment performs two-stage validation: first, schema validation; second, geometry feasibility checked by the Flat-Folder solver against Maekawa’s and Kawasaki’se theorems to ensure at least one feasible solution exists. If valid, the internal state is updated and new observations are rendered.

Evaluation Protocol

Improvements for AI systems

Here are specific improvements that can be made to AI systems by leveraging the insights from the OrigamiBench paper, along with what those improved systems could accomplish:


The core improvement lies in moving AI models beyond pattern recognition toward developing a deep, causal understanding of physical constraints and sequential decision-making.

Specific improvements include:

Improvement in Causal Reasoning over Visual/Language Representations:

A model should be trained or fine-tuned to explicitly ground its language understanding (symbolic concepts) in the visual perception of geometric transformations. This involves creating intermediate representations that explicitly map visual features to physical constraints (e.g., crease patterns, layer ordering).

Improvement in Multi-Step Planning and Coherent Strategy Generation:

The system needs a mechanism to generate long-horizon, multi-step folding strategies that are physically valid from the start. Instead of treating each fold as an independent visual prediction, the model should learn sequential dependencies where the output of fold 'n' provides necessary geometric context for fold 'n+1'.

Improvement in Constraint Adherence and Physical Validity Checking:

Integrating a robust physical validity module (like the Flat-Folder solver mentioned) directly into the planning loop, rather than relying solely on visual similarity metrics, ensures that proposed actions respect fundamental laws of paper folding (Maekawa's and Kawasaki's theorems).

Improvement in Interactive Learning and Execution-Based Supervision:

Transitioning from an open-loop generate final output paradigm to a closed-loop, interactive environment where the model receives continuous feedback (validity/similarity) after every action. This allows for reinforcement learning or iterative refinement based on physical success rather than just terminal correctness.

The improved AI system, leveraging these advancements, could perform the following:

Capable of Autonomous Origami Synthesis: The system would be able to take a textual description of a complex 3D shape (or even an abstract target) and generate the complete, executable sequence of precise folding actions required to construct it from a blank sheet of paper.

Advanced Geometric Problem Solving: It could solve complex geometric puzzles that require sequential manipulation, such as tangram assembly or intricate spatial packing problems, by treating the physical constraints as hard rules rather than soft visual cues.

Robust Physical Simulation and Design: The system could act as a design assistant for physical products (like packaging or architectural models), ensuring that any proposed structure is not only aesthetically similar to a target but is also guaranteed to be physically constructible and stable under real-world folding constraints.

Enhanced Generalization Across Domains: By learning the causal mechanisms of geometric transformation, the system could potentially transfer its reasoning skills to other structured, constraint-based domains that require sequential planning, such as robotics manipulation tasks or complex circuit design where physical rules govern state transitions.

Abstract

Building AI systems that can plan, act, and create in the physical world requires more than pattern recognition. Such systems must reason about the generative mechanisms and constraints governing physical processes, using structured representations that connect observations, actions, and their effects. Yet, many existing benchmarks study these capabilities separately, focusing either on visual recognition or on abstract symbolic or programmatic reasoning. Origami provides a natural testbed that integrates these abilities: constructing shapes through folds requires visual perception, reasoning about geometric and physical constraints, and sequential planning, while remaining sufficiently structured for systematic evaluation. We introduce OrigamiBench, a benchmark for evaluating programmatic understanding of the mechanisms underlying origami synthesis through a high-level language of physically grounded fold actions. Experiments with modern vision-language models reveal that scaling model size alone does not reliably improve reasoning about physical transformations. Moreover, models struggle to ground programmatic information in visual observations, suggesting that visual and language representations remain weakly integrated.

Related papers