PhysMirror: Physics-Aware Mirror Object Generation

arXiv:2607.03470 · cs.CV · Submitted 2026-07-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "PhysMirror: Physics-Aware Mirror Object Generation".

Jane: PhysMirror introduces a novel, end-to-end physics-aware generation framework designed to synthesize photorealistic images with geometrically correct mirror reflections,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, to get into the details of "PhysMirror: Physics-Aware Mirror Object Generation," the title itself tells us right away that this work is about building a system that understands and enforces physics when generating images with mirrors. The authors are Xuan-Bach Mai, Duy-Phuc Nguyen, Quoc-Van Le, Tam V. Nguyen, Thanh-Toan Do, Huu Le, Duong-Van Nguyen, Minh-Triet Tran, and Trung-Nghia Le.

Jane: Those names sound like a solid team of researchers tackling this problem from different angles in three dee modeling and generative AI. The implication here is that they are trying to solve the fundamental issue where current models fail at reflecting objects correctly in a mirror setup.

Lu: They are aiming to move beyond treating reflections as simple 2D completion tasks by integrating explicit three dee spatial priors into the generation process, which is a big conceptual step for how we approach scene synthesis.

Meng: So, they're suggesting that instead of relying solely on the AI's learned visual patterns for reflections, they are building a structured environment first to provide mathematical correctness. That makes me wonder about the computational cost of lifting everything into three dee meshes and then back again.

Lalam: It’s about making sure the AI has a verifiable structure to work with, which should lead to much more reliable outputs for any application, not just image generation.

The paper's summary: Tom: The summary of "PhysMirror: Physics-Aware Mirror Object Generation" explains that the framework is an end-to-end pipeline that starts with a text prompt and ends with a photorealistic image that has physically correct mirror reflections. They achieve this by first using a text-to-three dee model to turn objects into meshes, then composing those meshes into an accurate three dee mirror scene, and finally extracting depth maps and segmentation maps to guide the image generation step.

Jane: It’s essentially a four-stage process where the AI builds the physical reality first—the three dee scene—and then uses specific pieces of information from that reality, like depth maps, to steer the final image generation process toward correctness.

Lu: The methodology is quite structured: Stage one is lifting objects into three dee meshes, Stage two is building a physically accurate mirror scene using a specific reflection equation for points on the object surface, and Stage three involves rendering spatial conditioning maps from that scene.

Meng: The paper mentions that the reflection of any point p on an object surface visible to the mirror is calculated using Equation (one), which is then applied directly to explicit three dee geometry to produce a "geometrically correct reflected mesh" whose winding order is reversed for correct surface normals. That level of geometric precision sounds very demanding computationally.

Lalam: That emphasis on extracting precise 2D conditioning elements like depth maps and segmentation maps seems crucial because those are the specific signals that help guide the downstream text-to-image model to produce that physically correct look.

The paper's improvements: Tom: One of the key suggested improvements is introducing a novel metric called the Mirror Consistency Score, or MCS, which they claim is a fully automated way to measure physical correctness without needing any ground-truth annotations or human input for evaluation. They also developed a new benchmark dataset called the Mirror Object Benchmark, MirrOB, to test this framework rigorously across different object complexities and scene compositions.

Jane: The MCS sounds fantastic because if we can't manually grade every reflection, having a score that analyzes projective consistency between objects and their reflections based on vanishing point convergence gives us an objective way to measure success.

Lu: The authors demonstrated that custom depth-conditioned LoRA and zero-shot conditioning frameworks can significantly outperform existing state-of-the-art text-to-image models like FLUX.one and SDXL in terms of geometric realism, achieving a Mirror Consistency Score of zero point seven four six on the MirrOB dataset.

Meng: The practical improvement here is that they’ve shown how these conditioning maps can be plugged directly into existing diffusion models using lightweight adapters, which suggests a path for integrating physics awareness without completely rebuilding the entire generation stack from scratch.

Lalam: This capability to condition downstream models with explicit spatial priors means we don't need perfect three dee rendering for every single image; we just need those reliable maps to guide the final synthesis, which makes training data creation much more feasible.

Conclusion: Tom: So, wrapping up "PhysMirror: Physics-Aware Mirror Object Generation," the main implication is that we have a framework that natively enforces projective geometry by explicitly modeling three dee space and extracting geometric priors to guide image synthesis, which significantly reduces the risk of hallucinated reflections in text-to-image models.

Jane: It’s about moving from models that might just look plausible to models that adhere to strict spatial rules, which is vital if we’re training AI for tasks where real-world physics matter, like robotics perception.

Lu: The development of the Mirror Object Benchmark dataset and the MCS metric provides a clear path for researchers to objectively compare how well different conditioning strategies handle multi-object mirror scenes without needing human labeling.

Meng: For practical engineering applications, this suggests that we can build synthetic training environments where the physical constraints are guaranteed by the generation pipeline itself, which makes building perception networks much safer and more reliable.

Lalam: Ultimately, PhysMirror shows how providing explicit spatial conditioning maps derived from a simulated three dee scene allows for powerful guidance in downstream diffusion models, which could improve the cultural standard of synthetic data quality across the board.

University of Science Ho Chi Minh City · University of Dayton · Monash University · VinFast

cs.CV

Submitted: 2026-07-03

Updated: 2026-09-30

Comments: Accepted to IROS 2026

Code: https://github.com/T2I-Mirror-Object/PhysMirror

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 91/100

The gist: PhysMirror introduces a novel, end-to-end physics-aware generation framework designed to synthesize photorealistic images with geometrically correct mirror reflections, addressing a critical failure

Key concepts

Text-to-3D Generative Model
This model takes a written description of an object from a text prompt and converts it into a detailed 3D mesh representation. This step is crucial because it gives the system the necessary physical shape information, ensuring that the objects in the generated scene have strong, accurate geometry before any reflection is calculated.
Mirror Reflection Transformation
This mathematical formula precisely calculates how a point on an object's surface should appear when reflected in a mirror. It uses explicit 3D spatial priors like the mirror's normal and distance to define the exact geometric transformation, ensuring reflections are mathematically correct rather than just simple 2D completions.
Spatial Conditioning Maps
These are derived visual maps—specifically depth and segmentation maps—that encode the physical layout of the 3D scene. The depth map shows distances, and the segmentation map identifies different elements like 'real object' or 'reflection,' providing robust spatial guidance to help image generators place things correctly.
Mirror Consistency Score (MCS)
This is a new metric used to automatically evaluate how physically correct an image is. It checks the projective consistency between objects and their reflections by looking at where lines converge, quantifying the tightness of this intersection cluster around the vanishing point to measure overall physical realism.

Terminology

Summary

PhysMirror introduces a novel, end-to-end physics-aware generation framework designed to synthesize photorealistic images with geometrically correct mirror reflections, addressing a critical failure point in current text-to-image diffusion models used for generating synthetic training data. This method overcomes the limitations of existing approaches that treat reflections as simple 2D completion problems by natively enforcing projective geometry through explicit 3D spatial priors. By automatically lifting prompted objects into 3D meshes and constructing a mathematically exact mirror scene, PhysMirror extracts precise 2D conditioning elements—such as depth maps and segmentation maps—that serve as robust guiding signals for downstream diffusion models, thereby ensuring physically correct mirror reflections.

PhysMirror Pipeline Overview

The proposed method is an end-to-end pipeline that transitions from a text prompt to a physically grounded 3D scene, then projects back to 2D spatial conditioning maps to guide photorealistic image generation. This process is structured into four distinct stages:

  1. Lifting prompted objects into 3D meshes via a text-to-3D generative model.

  2. Composing a physically accurate mirror scene within a simulated 3D environment.

  3. Rendering conditioning maps, such as depth and segmentation maps, from optimized viewpoints using a standard 3D rasterization engine.

  4. Guiding a text-to-image generative model with these extracted spatial priors through frameworks like zero-shot conditioning or custom adapters.

3D Mesh Generation and Scene Composition

The foundation of the pipeline is establishing an explicit 3D representation for each object in the scene. Given an input text prompt, primary objects are parsed and lifted into 3D meshes using a text-to-3D model to ensure strong shape fidelity. Once meshes are generated, they are assembled into a physically accurate mirror scene. The mirror is defined as a plane with unit normal 'n' at signed distance 'd' from the origin, and the reflection of any point 'p' on an object surface visible to the mirror is computed using Eq. (1):

p′ = p − 2 (n · p − d) n. This transformation is applied directly to explicit 3D geometry, producing a geometrically correct reflected mesh whose winding order is reversed for correct surface normals.

Spatial Conditioning Extraction

With the 3D scene fully assembled, a virtual camera is positioned at a specific coordinate (Xc, Yc, Zc) where Zc > 0. A key design choice is that the camera's azimuth angle is randomly sampled to introduce a deliberate horizontal offset, increasing viewpoint diversity while ensuring the reflection remains within the mirror’s boundaries. This renders two crucial spatial conditioning maps:

  1. Depth Map: This map explicitly encodes the layered distance structure, where the real object is closest, followed by the mirror surface, and finally, the reflected object appears furthest behind.

  2. Segmentation Map: A per-region segmentation map assigns a distinct color to elements such as the real object, its reflection, mirror frame, floor, etc., with each region paired with a short descriptive text prompt.

Conditioning-Guided Generation and Evaluation

The extracted 3D spatial priors are integrated into a text-to-image generative model using conditioning strategies. The primary approach utilizes the pretrained OminiControl framework [10], which injects the rendered depth map as a lightweight plug-in adapter into FLUX.1-dev. Alternative strategies investigated include segmentation-conditioned generation through Seg2Any [11] and a custom depth-conditioned Low-Rank Adaptation (LoRA) fine-tuned on the SynMirror dataset [4]. To address the lack of automated evaluation tools, PhysMirror introduces the Mirror Consistency Score (MCS), a reference-free, fully automated metric that quantifies physical correctness by analyzing projective consistency between objects and their reflected counterparts based on vanishing point convergence.

Benchmark Dataset and Quantitative Results

To rigorously test the framework, a new benchmark dataset, the Mirror Object Benchmark (MirrOB), was constructed. This dataset consists of 360 structured prompts across three difficulty levels: single-object, two-object, and three-object scenes. The MCS metric evaluates correctness by first isolating objects and reflections using Grounding DINO [23] and SAM [24], extracting dense features via DINOv2 [25], enforcing a strict bijective mapping using a Mutual Nearest Neighbor (MNN) constraint, and finally measuring the tightness of this intersection cluster around the effective vanishing point to calculate MCS. Experimental results on MirrOB demonstrate that PhysMirror outperforms state-of-the-art baselines in geometric realism, achieving an overall Mirror Consistency Score of 0.746 and a CLIP Score of 28.914, while maintaining strong semantic text-image alignment.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems by implementing the PhysMirror framework, and what those improved systems will be able to do:


  1. Replacement of hallucinated or geometrically inconsistent reflections in text-to-image generation with physically accurate ones.

  2. Enforcement of strict projective geometry constraints during image synthesis, ensuring that reflected objects maintain mathematically correct spatial relationships (e.g., correct perspective and depth layering) relative to the mirror plane, eliminating structural hallucinations like missing reflections or mismatched object poses.

  3. Creation of a robust pipeline for generating synthetic training data for embodied AI and robotic perception by providing geometrically grounded images that adhere to real-world physics constraints, leading to safer and more reliable perception networks.

  4. Introduction of an automated, reference-free evaluation metric (Mirror Consistency Score - MCS) that quantitatively measures the physical correctness of mirror reflections using dense feature matching and vanishing point convergence analysis, allowing for rigorous benchmarking of physics-aware generative models without human intervention.

  5. Ability to generate diverse, physically consistent scenes containing multiple objects in front of a mirror (single, two-, or three-object scenes) where objects are correctly grounded on a floor plane and positioned at strict distances from the mirror, addressing the lack of spatial coherence seen in current state-of-the-art models.

  6. Enabling downstream diffusion models to be guided by explicit 2D spatial conditioning maps (depth maps and segmentation masks) extracted from a simulated 3D environment, allowing for zero-shot or lightweight conditioning strategies that bridge the gap between semantic creativity and physical laws.

  7. Development of specialized training data (MirrOB dataset) designed to test model robustness in complex multi-object mirror scenarios, specifically evaluating how models handle perspective, object details, and occlusion in physically realistic reflections.

Abstract

Synthesizing physically accurate mirror reflections remains a fundamental challenge for modern text-to-image diffusion models, which are increasingly critical for generating synthetic training data for embodied AI and robotic perception. These models typically struggle with strict geometric constraints, leading to hallucinations that degrade the utility of the synthetic data. To address this, we introduce a novel, end-to-end physics-aware generation framework namely PhysMirror that natively enforces projective geometry through explicit 3D spatial priors. Our method automatically lifts prompted objects into 3D meshes and constructs a lightweight, mathematically exact mirror scene within a simulated environment. By rendering this explicit 3D scene, we extract precise 2D conditioning elements, such as depth maps and segmentation maps, that serve as robust guiding signals for downstream diffusion models, guiding them to generate images with physically correct mirror reflections. Moreover, we introduce Mirror Consistency Score (MCS), reference-free, fully automated metric that quantifies physical correctness using dense feature matching and vanishing point convergence. Experimental results on our newly constructed MirrOB dataset demonstrate that our approach outperforms state-of-the-art baselines in reflection accuracy and physical realism, while maintaining strong text-to-image semantic alignment, providing a reliable pipeline for embodied AI data generation. The source code is released at https://duyphuc0701.github.io/PhysMirror.

Sources

Related papers