MatLat: Material Latent Space for PBR Texture Generation

arXiv:2512.17302 · cs.CV · Submitted 2025-12-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "MatLat: Material Latent Space for PBR Texture Generation".

Jane: The gist The authors propose MATLAT,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: The paper explains that standard methods like Score Distillation Sampling often produce textures with saturation artifacts because they don't handle PBR channels well.

Jane: They are proposing this MATLAT framework to solve that by leveraging the priors from those pretrained image diffusion models instead of starting from zero.

Tom: The main claim is that they fine-tune the pretrained VAE so that incorporating new material channels, like roughness and metallic, happens with minimal deviation from the original latent distribution.

Lu: They do this through a two-stage pipeline where the first stage adapts the latent space in MATVAE, and then the second stage uses a diffusion model to generate multi-view material images in that adapted space.

Meng: That second part is interesting because they address preserving multi-view consistency, which is usually tricky when you're working with different viewpoints.

Lalam: They handle that consistency by using something called correspondence-aware attention in the diffusion model, but they also add locality regularization to keep the latent and image pixels spatially aligned.

Conclusion: Tom: So we've looked at how this MatLat framework uses a two-stage approach to adapt an image diffusion model for generating PBR textures, focusing on learning that material latent space.

Jane: The authors are essentially showing how you can take existing knowledge from general images and tailor it specifically for the material properties of things you see in three dee models <ref:2512.17302#pg1>.

Tom: The authors found that this method outperforms methods trained from scratch because those lack the necessary PBR supervision, and it beats SDS methods which often show those saturation issues we talked about.

Lu: The paper suggests that by making sure the latent-to-image mapping maintains spatial locality through their regularization techniques, you get much better multi-view consistency for those material images.

Meng: For someone working on three dee assets, this means they can generate textures quickly and with higher fidelity than previous approaches, which is a practical gain <ref:2512.17302#pg1>.

Lalam: Culturally speaking, this kind of progress in generative modeling means we can create incredibly realistic digital content much more easily across all sorts of objects.

KAIST

cs.CV

Submitted: 2025-12-19

Updated: 2026-10-08

Project page: https://matlat-proj.github.io

Importance score: 85/100

The gist: The gist The authors propose MATLAT, a generative framework that learns a material latent space to produce high-quality PBR textures by leveraging pretrained image diffusion models and addressing

Key concepts

MATVAE
This is a modified Variational Autoencoder that learns a material latent space. It extends the original latent space to include PBR channels like roughness and metallic maps. It uses residual prediction and KL regularization to ensure the learned latent distribution stays close to the strong prior learned from pretrained models.
Latent-Space Adaptation
This process modifies the pretrained model's latent space so it can handle new information, specifically PBR channels. It involves injecting roughness and metallic details via a residual encoder and using regularization to keep the adapted space consistent with the original distribution, effectively bridging the gap between general images and material maps.
Correspondence-Aware Attention (CAA)
This technique is used within the diffusion model to ensure multi-view consistency. It restricts attention calculations only to points that correspond across different views of a 3D mesh. This explicit alignment strengthens the generation of PBR material images from multiple perspectives.
Locality Regularization (Llocal)
This regularization ensures spatial coherence in the latent space by enforcing patch-wise reconstruction during training. It forces each image pixel to be reconstructed primarily from aligned latent tokens, guaranteeing that the generated texture maintains strong spatial locality across different views.

Terminology

Summary

The gist The authors propose MATLAT, a generative framework that learns a material latent space to produce high-quality PBR textures by leveraging pretrained image diffusion models and addressing domain gaps through latent-space adaptation and locality regularization.

Motivation

Recent works on PBR texture generation advance beyond baked appearances by predicting material maps that enable relightable assets, but progress remains constrained by the scarcity of large-scale, highquality datasets with PBR channels The authors propose a generative framework for producing highquality PBR textures on a given 3D mesh. A natural direction to overcome this challenge is to leverage pretrained image generative models, which provide strong priors learned from large-scale RGB image datasets. Techniques such as Score Distillation Sampling (SDS) have been explored for this purpose, but they often fail to produce highfidelity PBR material images and exhibit saturation artifacts.

MATLAT Framework

The proposed framework is a two-stage pipeline where the first stage learns the material latent space by fine-tuning a pretrained VAE into Material VAE (MATVAE), and the second stage fine-tunes a diffusion model to generate multi-view material images in this adapted latent space.

  1. The first component is a latent-space adaptation module in MATVAE that extends the pretrained latent space to incorporate additional PBR channels—roughness and metallic—while preserving the pretrained prior. This approach fine-tunes the pretrained encoder to effectively incorporate the additional channels of PBR material images while applying a distributional regularization that constrains deviations from the original latent distribution.

  2. The second component is locality regularization in MATVAE, which establishes latent–image spatial alignment. This alignment is subsequently leveraged by correspondence-aware attention (CAA) within the diffusion model.

MATVAE Details

MATVAE encodes PBR material images into a latent space aligned with the pretrained latent distribution via two main mechanisms.

  1. Residual Prediction: The authors introduce a learnable residual encoder Eres to inject the roughness and metallic information, predicting residual parameters (µres,σres) = Eres(x), which adjust the base latent distribution such that z ∼ N µbase + µres,σ2 base ⊙ σ2 res = q(zx).

  2. KL Regularization: A regularizer Lreg is introduced to penalize the divergence between the learned latent distribution and that of the pretrained model, constraining the learned latent distribution to remain close to the pretrained latent space.

Multi-View Consistency

To ensure multi-view consistency, two components are used in MATVAE.

  1. Correspondence-Aware Attention (CAA): CAA incorporates explicit point correspondences into the attention computation by restricting computation to the correspondence set C(u). This strengthens cross-view alignment, yielding multi-view consistent PBR material images.

  2. Locality Regularization (Llocal): To ensure latent–image spatial locality, Llocal enforces patch-wise reconstruction during MATVAE fine-tuning by ensuring each image pixel is reconstructed mainly from aligned latent tokens. This regularization replaces the reconstruction loss in Eq. 2 with Llocal.

MATLAT Training

The final MATLAT model is trained by fine-tuning a latent diffusion (velocity) model uθ on multi-view PBR latents Z obtained by rendering the mesh from N viewpoints and encoding each view with MATVAE. This training follows the Conditional Flow Matching (CFM) objective defined in Eq. 9.

Results

Experiments show that MATLAT outperforms baselines trained from scratch, which suffer from suboptimal quality due to limited PBR supervision, and SDS-based methods, which tend to produce oversaturated appearances and exhibit prohibitive runtimes. Notably, our method also outperforms multi-view diffusion baselines on most quantitative metrics, achieving new stateof-the-art performance. Ablation studies validate the effectiveness of each component of our framework, including the proposed latentspace adaptation in MATVAE, correspondence-aware attention, and locality regularization. Qualitative results show that combining residual prediction with Lreg effectively preserves the pretrained latent distribution and yields high-quality PBR textures.

Conclusion

The authors present MATLAT, a two-stage pipeline that leverages the priors of pretrained image diffusion models for PBR texture generation. Directly extending image diffusion models to PBR material images is challenging due to the domain gap introduced by roughness and metallic channels. The framework introduces effective latentspace adaptation, MATVAE, which incorporates PBR material images into the pretrained latent distribution, together with locality regularization that enhances multi-view consistency during correspondence-aware attention computation. Experimental results show that our MATLAT achieves superior performance in PBR texture generation.

Potential Negative Societal Impacts

The proposed framework could be misused to generate hyper-realistic synthetic 3D assets that facilitate misinformation or deceptive digital content. The paper concludes by presenting qualitative results demonstrating strong generalization across diverse object categories and material types. The authors report the runtime in seconds required to generate a PBR texture.

References

[1] Mishan Aliev, Dmitry Baranchuk, and Kirill Struminsky. Castex: Cascaded text-to-texture synthesis via explicit texture maps and physically-based shading. 2025. 2

[2] Raphael Bensadoun, Yanir Kleiman, Idan Azuri, Omri Harosh, Andrea Vedaldi, Natalia Neverova, and Oran Gafni. Meta 3d texturegen: Fast and consistent texture generation for 3d objects. arXiv preprint arXiv:2407.02430, 2024. 13

[5] Dave Zhenyu Chen, Yawar Siddiqui, Hsin-Ying Lee, Sergey Tulyakov, and Matthias Nießner. Text2tex: Text-driven texture synthesis via diffusion models. In ICCV, 2023. 1

[7] Zilong Chen, Yikai Wang, Wenqiang Sun, Feng Wang, Yiwen Chen, and Huaping Liu. Meshgen: Generating pbr textured mesh with render-enhanced auto-encoder and generative data augmentation. In CVPR, 2025. 1

[8] Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl Vondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. Objaverse-xl: A universe of 10m+ 3d objects. In NeurIPS, 2023. 1

[10] Kangle Deng, Timothy Omernick, Alexander Weiss, Deva Ramanan, Jun-Yan Zhu, Tinghui Zhou, and Maneesh Agrawala. Flashtex: Fast relightable mesh texturing with lightcontrolnet. In ECCV, 2024. 2

[17] Zebin He, Mingxin Yang, Shuhui Yang, Yixuan Tang, Tao Wang, Kaihao Zhang, Guanying Chen, Yuhong Liu, Jie Jiang, Chunchao Guo, and Wenhan Luo. Materialmvp: Illumination-invariant material generation via multi-view pbr diffusion. In ICCV, 2025. 2

[18] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NeurIPS, 2017. 6

[21] Xin Huang, Tengfei Wang, Ziwei Liu, and Qing Wang. Material anything: Generating materials for any 3d object via diffusion. In CVPR, 2024. 2

[27] Akshay Krishnan, Xinchen Yan, Vincent Casser, and Abhijit Kundu. Orchid: Image latent diffusion for joint appearance and geometry generation. In ICCV, 2025. 2

[30] Zhibing Li, Tong Wu, Jing Tan, Mengchen Zhang, Jiaqi Wang, and Dahua Lin. Idarb: Intrinsic decomposition for arbitrary number of input views and illuminations. In ICLR, 2025. 3

[31] Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In CVPR, 2023.

Improvements for AI systems

  1. Bold header: Material Latent Space Adaptation (MATVAE)

The improved system can effectively incorporate additional PBR channels (roughness and metallic) into a single latent code while preserving the pretrained prior, addressing the problem where the encoded PBR materials and the pretrained latent space [experience] a substantial domain gap.

  1. Bold header: Locality Regularization (Llocal)

This component enforces spatial alignment by ensuring each image pixel is reconstructed mainly from its spatially local latent tokens, which prevents information exchange between geometrically unrelated tokens and enforces spatial locality between latent tokens and image pixels.

  1. Bold header: Multi-View Consistency Enhancement

The system can generate PBR textures that exhibit superior view alignment, achieving the highest cPSNR without performance degradation by combining Correspondence-Aware Attention (CAA) with Locality Regularization to ensure multi-view consistent PBR material images.

Sources

Related papers