Embedding Physical Reasoning into Diffusion-Based Shadow Generation Under the Sun and Sky

arXiv:2512.06174 · cs.CV · Submitted 2025-12-05 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Embedding Physical Reasoning into Diffusion-Based Shadow Generation Under the Sun and Sky".

Tom: Generating realistic shadows for inserted objects requires reasoning about scene geometry and illumination, which existing methods often fail to do purely in image space.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, we're talking about "Embedding Physical Reasoning into Diffusion-Based Shadow Generation Under the Sun and Sky" by Hu, Xu, Dave, Samaras, and Le from Stony Brook University. It’s a title that tells you right away they are integrating physical understanding directly into their diffusion shadow generation process.

Jane: That title really captures the essence of what they're doing—taking something that is usually learned implicitly and explicitly embedding reasoning about the physics of shadow formation into it, which is quite a detailed approach for an image generation task.

Lu: The authors are tackling the core difficulty: generating realistic shadows in scenes where you don't have perfect reference information, moving away from just looking at pixels to understanding the three dee scene structure. This moves the problem from pure pattern matching to geometric estimation.

Meng: I'm curious about their approach regarding how they manage that physics reasoning part; it sounds like a complex step before the actual diffusion happens. Does this rely heavily on prior knowledge or is it something they are trying to recover from scratch?

Lalam: The fact that they are recovering approximate scene geometry and estimating a dominant light direction using tools like MoGe-two forty-two suggests they are building an explicit spatial anchor, which could be a really important cultural asset for how we train future multimodal models.

The paper's summary: Tom: Looking at the summary of "Embedding Physical Reasoning into Diffusion-Based Shadow Generation Under the Sun and Sky," they outline a three-stage process. First, they use MoGe-two to get an approximate three dee point map and estimate a light direction, then they use those to compute a physics-based shadow estimate via geometric ray reasoning twenty-seven thirty-four.

Jane: That initial step sounds like it’s establishing the physical relationship between the object's shape and the incoming light source before anything else is generated, which makes sense for getting things right.

Lu: The summary highlights that this coarse shadow estimate acts as a spatial anchor for generation, encoding the object-receiver–illumination relationship before any learning takes place. This pre-computation step is key because it gives the generator something concrete to work with geometrically.

Meng: But then they move to refining this coarse estimate using a "mask corrector" to get a full-resolution auxiliary shadow mask, denoted as M˜f s, which is an interesting way to bridge that gap between rough estimation and fine detail.

Lalam: That refinement process sounds like it's where the heavy lifting of making the final shadow look sharp happens, and if they can handle that well, it could make our generated scenes look incredibly professional.

The paper's improvements: Tom: The authors highlight a significant improvement by showing their physics-grounded approach produces shadows that align more faithfully with the occluder geometry and scene lighting compared to methods like SGDiffusion twenty-two and GPSD fifty-one.

Jane: That comparison is telling; they claim better alignment even in scenes without reference background-object-shadow pairs, which tackles a major weakness in previous state-of-the-art techniques.

Lu: Their main contribution is explicitly introducing the framework that derives an initial shadow estimate from recovered monocular scene geometry and predicted illumination, providing that physically grounded spatial anchor for shadow placement. They also mention improving fidelity in top-down views under complex object geometry.

Meng: So, beyond just better alignment, does this mean a tangible reduction in the boundary errors when we look at the generated shadows? I'm interested in practical metrics for how much better it is than what we're currently seeing.

Lalam: Yes, because they specifically focus on producing more realistic and physically consistent shadows that respect the occluder’s shape, which should lead to a lower boundary error rate overall.

Conclusion: Tom: So, to wrap up the "Embedding Physical Reasoning into Diffusion-Based Shadow Generation Under the Sun and Sky" paper, they've shown how combining geometric recovery with confidence scores allows them to condition a frozen Stable Diffusion generator effectively through spatial and lighting guidance pathways.

Jane: It really boils down to using learned quality scores for both the lighting direction and the shadow mask so that uncertain cues get down-weighted during synthesis, which makes the whole generation process more reliable.

Lu: This ability to predict confidence scores for both light and mask cues, qlight and qmask, is what lets them use two complementary conditioning pathways in the diffusion model. It’s a sophisticated way to guide the AI synthesis based on its own uncertainty.

Meng: From my side, seeing that they can handle ambiguous scenes without reference pairs is very encouraging; it means we could deploy this framework in scenarios where we don't have perfect ground truth data for scene geometry or lighting, which is a big deal for real-world applications.

Lalam: I think the implication here is that AI can start reasoning about the physical world with much higher fidelity, and this work lays a foundation for more reliable creative content generation across the board.

Shilin Hu, Jingyi Xu, Akshat Dave, Dimitris Samaras, Hieu Le

Stony Brook University

cs.CV

Submitted: 2025-12-05

Updated: 2026-09-29

Comments: Project page: https://shilin21.github.io/physical_generation/

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Generating realistic shadows for inserted objects requires reasoning about scene geometry and illumination, which existing methods often fail to do purely in image space.

Key concepts

Physics-based Prior
This step estimates the shadow by checking if any object pixel lies between a receiver pixel and the estimated light source. It uses geometric reasoning to create an initial, physically determined shadow shape before any learning occurs.
Quality Estimator
This network predicts confidence scores (qlight and qmask) for the predicted light direction and shadow mask. These scores are vital because they allow the system to automatically down-weight unreliable predictions during image generation.
Spatial Control Branch
This pathway conditions the diffusion model on local, dense guidance derived from the refined shadow mask. The mask is scaled by its confidence score (qmask), meaning low-confidence masks have minimal influence on the final shadow detail.
Cross-Attention Tokens
This pathway uses estimated global lighting information to condition the diffusion model. The light direction estimate is modulated by its confidence score (qlight) using FiLM, providing scene-level orientation cues.

Terminology

Summary

Generating realistic shadows for inserted objects requires reasoning about scene geometry and illumination, which existing methods often fail to do purely in image space. This work introduces a framework that grounds shadow generation in the physics of shadow formation by recovering approximate scene geometry and estimating a dominant light direction to derive a physics-grounded shadow estimate via geometric reasoning.

How it works

The proposed framework operates in three main stages:

  1. A light predictor estimates a 3D light direction from the hint stack, which is constructed from the input image, object mask, and an approximate 3D point map recovered using MoGe-2 [42]. Using this estimated light direction along with the point map and object mask, rays are cast to obtain a coarse physics-based soft shadow prior. This prior is then refined by a mask corrector into a full-resolution auxiliary shadow mask, denoted as M˜f s.

  2. A quality estimator predicts two scalar confidence scores, qlight and qmask, measuring the reliability of the predicted light direction and shadow mask, respectively. These scores are crucial for regulating generation: reliable predictions guide the synthesis, while uncertain ones are automatically down-weighted.

  3. A frozen Stable Diffusion generator is conditioned on these signals through two complementary pathways: a spatial control branch for dense local guidance (using M˜f s scaled by qmask) and cross-attention tokens for global scene-level lighting guidance (using ˆl modulated by qlight).

Key Components and Physics Grounding

The core of the method is its grounding in the physics of shadow formation, where given a 3D object and a light direction, light rays from the source are occluded by the object, casting a shadow whose shape and placement are physically determined on receiver surfaces. The framework explicitly links occluder shape and light direction to a geometry-consistent shadow estimate, providing an explicit spatial anchor for generation before any learning takes place.

The shadow mask estimation involves two steps:

Physics-based prior:

The prior is computed by checking for each receiver pixel whether any object pixel lies between it and the light source. This is quantified by measuring how well the unit vector connecting object and receiver points aligns with the light direction, yielding a score s(o,r) which determines the prior score: prior(r) = max over o in O, s(o,r).

Coarse-to-fine refinement:

The soft shadow prior is refined to capture fine details. This involves two stages:

  1. At the coarse stage (64×64), the soft prior is concatenated with downsampled inputs and passed through a lightweight network to produce coarse logits, upsampled to full resolution as Lc.

  2. At the fine stage, a residual correction ∆ is predicted from Lc, H, and a broadcast light map, which is added in logit space to yield the final mask: tilde M fs = sigma (Lc + Delta).

Quality-Aware Conditioning for Ambiguous Lighting

To handle scenes where illumination cannot be uniquely inferred from a single image—a common issue—the framework models the reliability of its predictions. The quality network predicts qlight and qmask. This is achieved by:

Spatial Inputs:

The quality network takes the downsampled hint stack (H64), the predicted mask (M˜f s,64), and the soft shadow prior as spatial inputs.

Lighting Head:

This head combines a pooled descriptor with lighting inputs to predict qlight. The lighting input is derived from a compact intermediate feature from the light predictor.

Mask Head:

This head combines the pooled descriptor with ˆl and a set of lightweight mask geometry cues (e.g., mask-object centroid offset, mask elongation, and disagreement between M˜f s,64 and the soft prior).

The training for this network uses a combination of a regression loss (Smooth-l1) and a pairwise ranking loss that enforces correct relative ordering across samples: the ranking term provides the primary supervision signal.

Diffusion-Based Refinement

The final shadow synthesis is performed by conditioning a frozen Stable Diffusion U-Net through two distinct pathways, each gated by its corresponding confidence score.

Dense spatial conditioning:

The spatial structure is communicated via a ControlNet branch that takes the hint stack augmented with the predicted shadow mask as an additional channel. Crucially, the mask is scaled by qmask, so a low-confidence mask contributes weakly while the remaining hint channels are unaffected.

Scene-level lighting conditioning:

The estimated light direction provides global orientation cues. A lightweight encoder processes the image and object mask into a feature map, which is modulated by ˆl via FiLM with modulation strength gated by qlight.

Improvements for AI systems

Here are specific improvements that can be made to AI systems by implementing the methodology described in this paper, along with a detailed description of what those improved systems can achieve:


The core contribution of this paper is shifting shadow generation from implicit, purely image-space learning to an explicit, physics-grounded reasoning framework. The improvements focus on grounding 2D generative models in 3D scene geometry and illumination.

Here are the specific improvements and their resulting capabilities:

  1. A shift from purely learned shadow synthesis to a hybrid system that integrates three distinct modules:

  2. An explicit, physics-based spatial anchor derived from recovered monocular scene geometry (using MoGe-2).

  3. A quality-gated diffusion refinement process that uses predicted confidence scores for lighting and shadow mask as dynamic weights on the generation process.

The resulting improved AI system can perform the following specific tasks:

  1. Generating highly realistic and geometrically consistent shadows in complex scenes, specifically when background reference pairs are scarce (BOS-free settings).

  2. Achieving superior shadow localization accuracy compared to state-of-the-art methods by anchoring shadows to recovered 3D point maps and estimated light directions.

  3. Robustly handling ambiguous or weakly lit scenes by using learned confidence scores to dynamically down-weight unreliable cues (e.g., lighting direction or mask prediction), preventing the hallucination of physically implausible shadows.

  4. Producing high-fidelity shadow masks that accurately respect the occluder's shape and its interaction with scene illumination, leading to a reduction in Shadow Region BER (Boundary Error Rate).

  5. Improving overall image quality and fidelity by ensuring that the synthesized shadows align faithfully with the underlying geometry, resulting in lower Root Mean Square Error (RMSE) in both global and local shadow regions.

In essence, this framework allows AI systems to move beyond merely painting shadows to reasoning about how light interacts with a 3D structure, leading to more physically plausible and reliable scene understanding tasks.

Abstract

Generating realistic cast shadows for inserted foreground objects requires reasoning about scene geometry and illumination. However, most learning-based approaches treat shadow generation as an image translation problem and capture these physical relationships only implicitly. This often results in misaligned or implausible shadows. Motivated by the physics of shadow formation, we introduce explicit geometric guidance for outdoor shadow generation. Given a composite image and a foreground object mask, we recover approximate scene geometry and estimate a dominant light direction to derive a coarse shadow estimate via geometric reasoning. While coarse, this estimate provides a spatial anchor for shadow placement. Because illumination cannot always be uniquely inferred from a single image, we predict confidence scores for both lighting and shadow cues and use them to regulate their influence during generation. These cues (shadow mask, light direction, and their confidence scores) condition a diffusion-based generator that refines the estimate into a realistic shadow. Experiments on DESOBAV2 show substantially improved shadow-region fidelity and localization, with an overall 23% lower shadow-region RMSE and 30% lower shadow-mask BER than the prior state-of-the-art method.

Sources

Related papers