MatLoom: Layered Text-to-Material Generation in a Compact Program Space

arXiv:2609.40322 · cs.CV, cs.AI, cs.CL, cs.MM · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "MatLoom: Layered Text-to-Material Generation in a Compact Program Space".

Jane: Compact, layer-oriented programs offer an effective output space for pretrained language models by combining text-to-material fidelity with explicit authoring structure.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, what’s the main idea behind this MatLoom paper? Basically, they introduce this compact language that represents a material as a stack of alpha-masked layers. This design lets each layer assign physically based rendering channels through expressions over a two-dimensional domain.

Jane: That sounds incredibly organized; it means they separate the editable source from the final sampled material maps, which is something we’ve been trying to achieve in generative systems for a while.

Lu: The paper claims that this structure makes dependencies between different patterns, color, and relief explicit within the program itself.

Meng: So it's not just a flat image; it's a structured definition of how the surface should look based on layered components and shared spatial expressions.

Lalam: Exactly, and they emphasize that named spatial fields can be shared across masks, colors, roughness, and height, which really highlights how these different material properties are coupled.

Conclusion: Tom: Looking at the title, "MatLoom: Layered Text-to-Material Generation in a Compact Program Space," it really sums up the core contribution of this work. The authors are showing that we can use a compact program space to generate material outputs that have an explicit authoring structure built right into them.

Jane: And what this means in simpler terms is that instead of just getting a picture, you get the instructions for building the picture, and those instructions are preserved alongside the final result.

Lu: The implication here is that pretrained language models can act as designers who write and repair these programs, giving us more control over the output’s physical properties.

Meng: It suggests a way to get more predictable material generation when you need specific structural features defined by the input prompt, rather than relying solely on the model's internal interpretation of texture.

Lalam: I think this points toward a future where generative AI isn't just producing pretty pictures, but is actively creating assets that are inherently structured and editable from their own construction logic.

Tom: So we’re looking at how this MatLoom framework lets the model generate materials with explicit construction rules, which is really impressive given the complexity of physical surfaces.

Jane: It really shows a path toward making generative AI outputs more predictable because you're defining the underlying rules up front.

Lu: I think the real impact could be in design industries where controlling material properties precisely is crucial for manufacturing or architectural visualization.

Meng: For practical use, having that explicit authoring structure means we can inspect and edit those construction rules, which is a huge win for iterative development.

Lalam: Ultimately, this work suggests that the next generation of generative models could be highly effective at producing complex physical assets because they understand the material's underlying composition as part of their generation process.

Anson Y. Lam, Shuqing Li*, Michael R. Lyu

Department of Computer Science and Engineering, The Chinese University of Hong Kong

cs.CV, cs.AI, cs.CL, cs.MM

Submitted: 2026-09-30

Updated: 2026-09-30

Importance score: 89/100

The gist: Compact, layer-oriented programs offer an effective output space for pretrained language models by combining text-to-material fidelity with explicit authoring structure.

Key concepts

MATLOOM Representation
MATLOOM represents a material as a stack of alpha-masked layers. Each layer assigns physically based rendering (PBR) channels using expressions over a 2D domain. This design makes dependencies explicit, allowing shared spatial fields to define coverage and surface properties across different maps.
Program Structure
A program consists of three parts: a sampling window (View), named expressions (Define), and bottom-to-top layers (Material). Expressions are functions over the XY plane that can be complex, enabling precise control over how layer values resolve at every position, ensuring edits in one field don't accidentally break others.
Synthesis Procedure
The synthesis involves a three-step process: first, a language model writes and repairs the program; second, a critic revises the design based on previews and statistics; and finally, a text–image scorer selects candidates. This separates material design from final realization for better control over the generation process.

Terminology

Summary

Compact, layer-oriented programs offer an effective output space for pretrained language models by combining text-to-material fidelity with explicit authoring structure.

The gist: Compact executable programs thus offer a way to generate prompt-aligned materials while retaining their construction as part of the asset.

MATLOOM Representation and Structure

MATLOOM represents a material as a stack of alpha-masked layers, where each layer assigns PBR channels through expressions over a two-dimensional domain. This design makes dependencies explicit, allowing shared spatial expressions to define coverage and physically based rendering (PBR) channels. Named spatial fields can be shared across masks, color, roughness, and height. The representation separates editable source from sampled material maps, where layers group coverage with surface properties while named fields expose cross-channel dependencies.

Program Structure and Semantics

A program consists of an optional sampling window (View), named expressions (Define), and bottom-to-top layers (Material). Expressions are fields over the XY plane, which can be constants or complex expressions built from transforms, noise, periodic patterns, and shapes. The compositing semantics define how layer values resolve at each position: for scalars like base color, it uses a convex blend of layer values, while height takes the maximum finite height among positively covered layers. This structure ensures that an edit confined to one field preserves other exported maps unless it is part of a shared dependency.

Synthesis and Synthesis Procedure

The synthesis procedure organizes inference by separating design from realization. A pretrained language model writes and repairs a program using parser feedback, followed by a critic that inspects a fast preview, channel statistics, and source code where configured to revise the material design. Finally, a text–image scorer selects candidates from the revision trajectory and searches noise seeds while keeping each candidate’s remaining source fixed. This process involves:

  1. A language model writing and repairing a program using parser feedback.

  2. A critic revising the material design based on preview, statistics, and source code.

  3. A text–image scorer selecting candidates from the revision trajectory and searching noise seeds while keeping each candidate’s remaining source fixed.

Evaluation and Empirical Results

MATLOOM was evaluated on a curated benchmark of 141 prompts with six language-model backbones against three diffusion baselines. The best configuration achieved higher mean scores than three diffusion baselines on all four flat-layout prompt-alignment metrics. In a blind four-way comparison involving 30 participants and 20 prompts, MATLOOM receives 59.2% of choices, compared with 19.3% for the most-preferred baseline. The strongest configuration, gemini-3.6-flash, had the highest mean on all four alignment metrics in both layouts (Table 1).

Program Control and Editability

The system retains construction as part of the asset through explicit authoring structure. Edits are inspectable; for instance, changing a shared mortar parameter updates both grout exposure and tile-edge relief simultaneously. The representation supports one-case edits that test these dependencies, demonstrating inspectable control for one material, not an editing success rate. Furthermore, inlining named definitions preserves evaluation semantics on tested programs without altering the exported channel maps.

Latency and Component Diagnostics

The flagship’s median run takes 4.6 minutes, with a median of 68% of time spent in polishing (Stage III seed search). The seed search stage involves rendering 1001 candidates at 256×256 previews in a median of 70 seconds. Exploratory component comparisons motivated by Table 9 show that the render-only critic has the largest observed flat judge delta (+8.50), while the render-plus-statistics critic has the largest flat BLIPScore delta (+6.96). These diagnostics motivate controlled tests but do not establish isolated component effects.

User Study Preferences

A blind online preference study over 20 prompts showed that MATLOOM won 59.2% of forced choices with a mean rating of 4.98, ahead of StableMaterials at 19.3%. The preference held descriptively for participants with no image-creation experience (60.0% MATLOOM win rate) and those with some experience (58.8%). This study measures prompt-conditioned appearance preference on this subset, not editing utility or physical accuracy.

Limitations

The fixed primitive and channel vocabulary constrain fine microstructure, specific figurative motifs, participating media, and arbitrary reflectance models. The benchmark measures prompt alignment, not physical material accuracy or edit success. Matched generation and editing comparisons against other executable material forms are still needed to isolate the value of this representation from backbone, budget, and renderer effects.

Improvements for AI systems

Here are specific improvements to AI systems based on the MATLOOM framework:

  1. Improving Material Generation Fidelity with Explicit Construction Rules:

  2. Enabling Iterative, Structure-Aware Material Revision (Critique and Repair):

  3. Enhancing Prompt Alignment through Controlled Stochastic Search (Seed Search):

  4. Developing a Compact, Executable Authoring Language (MATLOOM) for Assets:


  1. Improving Material Generation Fidelity with Explicit Construction Rules:

The improved system can generate materials where the underlying rules (spatial patterns, coverage masks, and PBR channel assignments) are explicitly defined in the source code rather than implicitly learned by a diffusion model.

  • Specific Capability: The system will produce materials where dependencies between spatial layout (e.g., grout width), color, roughness, and relief are mathematically explicit through shared spatial expressions (e.g., using a single mask to drive both coverage and height).

  • Benefit: This ensures that the generated asset retains its construction logic. Designers can inspect, re-evaluate, and revise the material structure without waiting for a complete re-generation by a diffusion model.

  1. Enabling Iterative, Structure-Aware Material Revision (Critique and Repair):

The improved system will utilize an LLM as an author that produces programs which are then subjected to a critique stage using channel statistics and source code analysis to suggest specific, structure-preserving revisions.

  • Specific Capability: The system can perform parser-guided repair where the LLM receives feedback on mismatches (e.g., the grout is too narrow) and revises only the relevant parameters in the source program while preserving unchanged layers and noise seeds.

  • Benefit: This moves beyond simple text-to-image iteration to genuine procedural editing, allowing designers to make localized structural changes (like widening grout or changing a specific glaze color) with high confidence that the rest of the material structure remains intact.

  1. Enhancing Prompt Alignment through Controlled Stochastic Search (Seed Search):

The improved system will employ a systematic search procedure that treats the program as an executable object, allowing for targeted exploration of stochastic variations (noise seeds) around high-performing candidates.

  • Specific Capability: The system can perform seed search by keeping the core program structure fixed while exploring thousands of noise seed variants to find superior realizations for a given prompt. The selection process uses a quick score and oracle to guide this search, maximizing the chance of finding an output that satisfies both prompt alignment and structural constraints.

  • Benefit: This provides a mechanism to generate multiple high-quality candidates for a single prompt, allowing the system to select the best stochastic realization rather than relying on a single generation run.

  1. Developing a Compact, Executable Authoring Language (MATLOOM) for Assets:

The improved AI system will leverage MATLOOM as its primary interface for material authoring, bridging natural language requests with executable procedural code.

  • Specific Capability: The system will parse natural language prompts into the MATLOOM DSL (defining views, named expressions over 2D domains, and layered material stacks), which is then evaluated by a standalone interpreter to produce physical PBR maps.

  • Benefit: This creates a compact, reusable asset format where the construction rules are intrinsically part of the final file. It allows for generating prompt-aligned materials while retaining their explicit construction as part of the asset itself, enabling future editing and reuse.

Sources

Related papers