MaPa: Text-driven Photorealistic Material Painting for 3D Shapes
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "MaPa: Text-driven Photorealistic Material Painting for 3D Shapes".
Jane: This research introduces MaPa, a novel framework designed to generate photorealistic and editable materials for 3D meshes directly from textual descriptions.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We’re looking at the title, "MaPa: Text-driven Photorealistic Material Painting for three dee Shapes," which really tells us the core mission here is achieving photorealism through text input specifically for materials on meshes. The authors are Zhang, Peng, Xu, Yang, Chen, Xue, Shen, Bao, Hu and Zhou from Zhejiang University and Ant Group.
Jane: That title really sets the stage because it emphasizes "text-driven," meaning the entire material look comes from a description rather than hand-painting or manual texture mapping. It’s a significant step away from older techniques where you had to painstakingly create every texture map yourself.
Lu: I think the authors are smart because they aren't just trying to generate one final image; they are proposing a segment-wise procedural material graph representation, which is different from synthesizing a single texture map for the whole object at once. This allows for finer control later on.
Meng: That segmentation idea sounds promising but also complex when you think about how the AI needs to decide where one segment ends and another begins based only on a text prompt. What kind of input is needed to reliably get those segments defined?
Lalam: The segmentation aspect is key because it breaks the problem down into manageable pieces, which makes it more tractable for the diffusion model to handle locally. This modular approach feels like a much better way for an AI to learn how physical objects actually look.
The paper's summary: Tom: So, what MaPa actually does is decompose the input mesh into segments, uses a segment-controlled diffusion model to generate 2D images aligned with those parts, and then optimizes procedural material graphs for each group. Basically, it’s a four-step pipeline to get from text to usable materials on a three dee shape.
Jane: That sounds like the core mechanism: first you get the image alignment using segment-controlled generation, then you group those segments by color and class, and finally, you select and optimize the best material graph for each group using a differentiable rendering module. It’s quite detailed in its process.
Lu: The way they use the SAM-conditioned ControlNet to generate images that are explicitly aligned with mesh parts is a clever technique to ensure stability during the subsequent optimization phase, which addresses some of the instability issues seen in other distillation methods.
Meng: I see what you mean about stability, but I wonder if this whole chain—segmentation, diffusion generation, grouping—is computationally heavy enough for real-time applications or even quick prototyping on a standard workstation. What’s the practical runtime like?
Lalam: The iterative recovery step is interesting because it handles those areas where materials aren't assigned initially by using inpainting networks to fill in the gaps, which makes the overall system more robust to missing information during the process.
The paper's improvements: Tom: Regarding improvements, MaPa focuses heavily on creating materials that are inherently editable by users in modeling software. The authors hope that by generating segment-wise graphs instead of monolithic textures, users can swap out material for just one area without having to rebuild the whole map.
Jane: That flexibility is huge because it mirrors how real manufacturing works, where different parts of an object often have distinct materials. They are trying to mimic that consistency found in real-life manufactured objects, which they observed as a motivation for this structure.
Lu: The improvement here lies in moving away from per-point representations toward these procedural material graphs, which the paper claims gives substantial flexibility for downstream user modifications, unlike methods that create single texture maps.
Meng: I'm still focused on the practical side—the paper notes that without grouping segments first, the optimization time can increase significantly to about thirty-three minutes per shape instead of just seven minutes when they use material grouping. That difference in speed is a major practical win for any engineer trying to iterate quickly.
Lalam: That speed improvement due to material grouping really speaks to how essential that coherence step is; it shows that making intelligent decisions about which segments belong together saves a ton of computational effort later on.
Conclusion: Tom: So, summarizing the main points of "MaPa: Text-driven Photorealistic Material Painting for three dee Shapes," we have this framework that uses segment-controlled diffusion to generate aligned images, groups those segments intelligently, and optimizes procedural material graphs for high-quality, editable materials.
Jane: In essence, it moves us closer to designing objects where the appearance can be defined by a description and then easily modified piece by piece within three dee software. It’s a significant step in bridging the gap between text prompts and tangible three dee assets.
Lu: The implication for AI research is that leveraging pre-trained 2D diffusion models as a bridge to material graphs is a very effective way to connect high-level semantic understanding with complex procedural generation tasks. We should look at how this influences other generative pipelines.
Meng: For practical impact, the main takeaway is that if we can achieve faster iteration times while maintaining quality, this becomes much more viable for designers and rapid prototyping workflows in industries that rely on three dee assets.
Lalam: I see this advancing how AI can shape our cultural perception of design; when materials become text-driven and highly editable, the creative process shifts from painstaking manual work to intelligent high-level direction, which is really powerful for creative expression.
Tom: Fantastic summary, everyone. We’ve covered a lot about MaPa today. It seems like this paper really shows how structure—segmentation and grouping—is key to unlocking the potential of text-driven material generation. We’ll be hearing more on this topic soon, so keep your ears tuned!
Zhejiang University · Ant Group
cs.CV
Submitted: 2024-04-26
Updated: 2026-10-01
Comments: Corrected the spelling of the first author's name in the manuscript and metadata; no changes to the technical content
Project page: https://zhanghe3z.github.io/MaPa
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 87/100
The gist: This research introduces MaPa, a novel framework designed to generate photorealistic and editable materials for 3D meshes directly from textual descriptions.
Key concepts
- Segment-controlled Image Generation
- The process breaks the 3D mesh into pieces and generates 2D images for each piece. A special diffusion model, conditioned on these segment masks, creates high-quality RGB images that accurately match the geometry of specific parts of the mesh. This ensures the generated textures align correctly with where they should be applied.
- Material Grouping
- Segments are grouped together based on shared material types and similar colors. This step simplifies the process by reducing optimization time and improving material coherence. GPT-4v is used to classify materials, while color similarity is measured using a specific distance threshold in the CIE color space.
- Material Graph Optimization
- For each material group, the best existing procedural material graph is retrieved. This graph's parameters are then optimized to match the generated image using a differentiable rendering module. This converts the abstract graph into usable texture maps like Albedo and Normal maps, ensuring photorealistic results.
- Iterative Material Recovery
- If some areas lack assigned materials, an iterative process is used. The system selects adjacent viewpoints and uses inpainting networks to fill gaps. Missing material assignments are then inpainted using a SAM-conditioned network until the entire object has a complete material assignment.
Terminology
Summary
This research introduces MaPa, a novel framework designed to generate photorealistic and editable materials for 3D meshes directly from textual descriptions. The method addresses the need for an efficient way to design appearances for meshes by proposing a segment-wise procedural material graphs representation, leveraging pre-trained 2D diffusion models as a bridge between text and material graphs.
The gist
MaPa generates photorealistic and high-resolution materials for meshes from textual descriptions by decomposing the shape into segments, using a segment-controlled diffusion model to synthesize aligned 2D images, and then optimizing segment-wise procedural material graphs through a differentiable rendering module.
How it works
The proposed framework consists of four main steps:
-
Segment-controlled image generation: The input mesh is oversegmented into segments, projected onto various viewpoints to produce 2D segmentation masks. A segment-controlled diffusion model, specifically the SAM-conditioned ControlNet, is then used to generate an RGB image from these 2D segments. This approach aims to create 2D images that
more accurately align with parts of the mesh,
enhancing stability for subsequent optimization. -
Material grouping: To improve coherence and save optimization time, segments are merged into groups based on shared material class and similar colors. Material classification is performed using GPT-4v, while color similarity is determined by comparing the median color distance between segments in the CIE color space against a threshold lambda (set to 2).
-
Material graph selection and optimization: For each material group, the most similar material graph is retrieved from a pre-built library using CLIP for zero-shot retrieval. The parameters of this retrieved material graph are then optimized to fit the generated image through a differentiable rendering module, DiffMat v2. This module uses physically-based microfacet BRDF and 2D spatially varying incoming lighting (predicted by InvRenderNet) to convert the procedural material graph into texture-space maps (Albedo, Normal, Roughness).
-
Iterative material recovery: To handle areas without assigned materials, an iterative approach is employed. The process selects adjacent viewpoints for the next iteration and inpainting networks are used to fill missing regions. Segments requiring material assignment are inpainted using the SAM-conditioned inpainting network, and the material graph generation process (steps 2 and 3) is repeated until all regions are assigned materials.
Key components of the pipeline include:
(a) Segment-controlled Image Generation:
(b) Material Grouping:
(c) Material Graph Selection and Optimization:
Downstream Editing:
The method provides a downstream editing module allowing users to modify materials via textual prompts through GPT-4 and predefined APIs, enabling high-level operations such as adding patterns.
Evaluation and Results:
Extensive experiments validate the framework's performance across different object categories (wood, metal, plastic, leather, fabric, stone, ceramic, rubber). Quantitative comparisons using FID and KID metrics show superior performance over baselines: Ours 88.3 / 0.037 / 4.33 / 3.35
for chairs and Ours 87.3 / 0.014 / 4.10 / 3.22
for the ABO dataset, outperforming TEXTure, Text2tex, and Fantasia3D in fidelity and quality ratings from user studies (Table 1). Qualitative comparisons demonstrate superior photorealism compared to baselines, which are prone to producing inconsistent textures
or failing with oversaturation. The framework also shows diversity in results when optimizing based on different diffusion model seeds and supports appearance transfer from reference images using IP-Adapter.
Limitations:
The paper notes limitations, including the Domain gap of generated images,
where the albedo estimation network sometimes fails due to the difference between the diffusion model's training data (natural images) and its synthetic training data. Additionally, for Complex objects with unobvious segments,
over-segmentation might occur, and the expressiveness of materials is currently limited by a lack of material graphs containing complex nodes. The optimization time without material grouping increases significantly to 33 minutes on average per shape compared to 7 minutes with grouping.
Future Directions:
The authors suggest exploring ways to train an amortized inference network that can directly predict the parameters of the material graph from the input image
for improved efficiency and considering fine-tuning diffusion models on synthetic images or employing more advanced albedo estimation networks to mitigate domain gap issues. The method also allows for editing materials in software like Substance Designer, offering flexibility beyond simple texture maps.
Acknowledgments:
The work was supported by NSFC, Ant Group, Information Technology Center and State Key Lab of CAD&CG at Zhejiang University. (This section is not part of the core methodology but is noted as required context.
Improvements for AI systems
Here are specific improvements that can be made to AI systems by leveraging the proposed MaPa framework:
-
The system can generate photorealistic, high-resolution, and editable 3D materials directly from natural language descriptions of objects (e.g.,
a wooden bedside table
). This moves beyond simple texture maps to procedural material graphs. -
The AI system can ensure material consistency across different parts of a mesh by segmenting the object first and then generating segment-specific material graphs, mimicking real-world manufacturing processes.
-
It can create materials that are highly editable by users in 3D software (like Substance Designer), allowing for complex modifications (e.g., adding luxury patterns via text prompts) without needing to recreate entire maps.
-
The system can perform appearance transfer, enabling the application of a reference image's material style onto an arbitrary 3D object based on a text prompt describing the desired look.
-
It can handle complex objects with non-obvious segment boundaries by intelligently grouping visually similar segments, leading to more coherent final appearances compared to methods that treat every segment independently.
-
The system can generate diverse material variations for the same object simply by generating multiple diffusion model outputs and optimizing the resulting materials, providing a rich set of aesthetic choices.
Abstract
This paper aims to generate materials for 3D meshes from text descriptions. Unlike existing methods that synthesize texture maps, we propose to generate segment-wise procedural material graphs as the appearance representation, which supports high-quality rendering and provides substantial flexibility in editing. Instead of relying on extensive paired data, i.e., 3D meshes with material graphs and corresponding text descriptions, to train a material graph generative model, we propose to leverage the pre-trained 2D diffusion model as a bridge to connect the text and material graphs. Specifically, our approach decomposes a shape into a set of segments and designs a segment-controlled diffusion model to synthesize 2D images that are aligned with mesh parts. Based on generated images, we initialize parameters of material graphs and fine-tune them through the differentiable rendering module to produce materials in accordance with the textual description. Extensive experiments demonstrate the superior performance of our framework in photorealism, resolution, and editability over existing methods. Project page: https://zju3dv.github.io/MaPa
Sources
- Text2Tex: Text-driven Texture Synthesis via Diffusion Models
- Deep convolutional filter banks for texture recognition and segmentation
- SIGNeRF: Scene Integrated Generation for Neural Radiance Fields
- Wavelet Convolutional Neural Networks for Texture Classification
- Adam: A Method for Stochastic Optimization
- Segment Anything
- Zero-1-to-3: Zero-shot One Image to 3D Object
- Wonder3D: Single Image to 3D using Cross-Domain Diffusion
- Latent-NeRF for Shape-Guided Generation of 3D Shapes and Textures
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
- PhotoShape: Photorealistic Materials for Large-Scale Shape Collections
- TEXTure: Text-Guided Texturing of 3D Shapes
- Denoising Diffusion Implicit Models
- TextureDreamer: Image-guided Texture Synthesis through Geometry-aware Diffusion
- Paint3D: Paint Anything 3D with Lighting-Less Texture Diffusion Models
- MateRobot: Material Recognition in Wearable Robotics for People with Visual Impairments
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models