LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting".
Jane: Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurate estimation of intrinsic scene properties such as materials, geometry, and illumination.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we're starting by looking at the title and who came up with it. This paper is titled "LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting," and the authors are Zixin Guo, Yehonathan Litman, Yifeng He, John Miller, Chuhan Chen, and Deva Ramanan. Jane It sounds like they're focusing on making video relighting more controllable while keeping the temporal consistency steady over long sequences.
Lu: From my perspective at Tsinghua University, this title immediately suggests a shift away from purely generative translation methods toward something that grounds the process in physics, which is really exciting. Meng I see it as an attempt to bridge the gap between high-level scene understanding and actual light transport simulation; that's where the real engineering challenge lies.
Lalam: Based on my analysis, I think the core idea here is using a hybrid approach to get both structural control and photorealism simultaneously, which could significantly impact how we use generative video tools for creative production.
Tom: Exactly. Jane So what's the main point they are trying to explain in that title? Lu Essentially, they are combining inverse rendering with diffusion refinement to achieve relighting that respects physical laws while maintaining consistency across time.
Meng: If I’m looking at the practical side, I’d be interested in how much of the control they actually retain once the system is trained. Tom Right, that's a big question for us.
The paper's summary: Jane: So, let's look at what the paper actually says about how this works. Tom The summary explains that LightCrafter reformulates video relighting as translating a proxy video by translating a Physically Based Rendering or PBR rendering under the target illumination conditions to the final target.
Lu: That framing is very interesting because it separates the problem into two distinct parts: baking the illumination targets into a PBR proxy and then using a diffusion model just for artifact correction on top of that. Tom So, they are using this PBR rendering as a structured way to capture most of the light interaction changes upfront.
Meng: That sounds like they’re trying to use physics to handle the heavy lifting of shadows and reflections, which saves the diffusion model from having to learn all those hard physical interactions from scratch. Jane It seems they are trying to get the temporal consistency baked into that PBR step because deterministic rendering handles that uniformly across frames.
Lalam: I think what’s most important here is how they manage the noise correction; by using a diffusion model just for refinement, they can focus on fixing those subtle, photorealistic details instead of trying to generate every pixel from scratch. Tom So it's a tiered approach: structure first with PBR, then detail correction with diffusion.
Jane: That makes sense in terms of pipeline design. Lu And the method they propose involves three main stages: first recovering scene intrinsics, second rendering a PBR video under the target illumination, and finally refining this PBR video into a photorealistic output via a video diffusion model.
The paper's improvements: Tom: Moving on to what actually makes LightCrafter better than what came before, the paper highlights several key improvements. Jane They focus heavily on making sure the relighting is controllable and consistent over long durations, which addresses a major weakness in previous methods.
Lu: The authors introduce a novel data curation pipeline specifically designed for artifact-matched supervision, where synthetic pairs expose the model to realistic rendering artifacts while still providing ground-truth PBR for supervision. Meng That’s smart; exposing the model to those realistic errors during training should help it correct those same errors when it encounters real footage later.
Tom: And they specifically mention that this approach enables consistent long-form relighting with far less drift than prior methods, which is a big win for any application involving continuous video sequences. Jane I think the ability to reuse a single trained refiner for different scene edits is another significant improvement because it simplifies the workflow immensely.
Lalam: If you think about the cultural impact of this, this consistency means we can produce long-form content that looks coherent without needing constant re-correction, which speeds up creative workflows substantially. Lu And they show generalization to tasks like object placement and material editing, meaning you don't have to retrain the whole system every time you want to change something in the scene.
Meng: From an engineering standpoint, being able to change things like adding new lights or modifying UV maps just by re-rendering the PBR video sounds incredibly efficient for deployment. Jane So, in short, they've focused on robustness and control by grounding the process in physical principles and smart data preparation.
Conclusion: Tom: Alright team, we've covered a lot about LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting. Jane To wrap up, the main implication here is that by grounding the relighting in physical principles through PBR rendering, they manage to get both structured control over scene properties and photorealistic results simultaneously.
Lu: This method shows a path where traditional physical simulations can be integrated with powerful generative AI models to tackle complex visual tasks like long-form video relighting. Meng I think the real impact is how this moves relighting from being a brittle, frame-by-frame correction task to something that operates on a more stable, physically consistent representation.
Lalam: For the culture, this means we can expect much more sophisticated AI tools for content creation that handle complex lighting scenes reliably without constant manual tweaking. Tom It's exciting stuff because they achieved state-of-the-art performance on existing benchmarks while solving the long-form consistency problem with overlap-fused temporal tiling.
Jane: So, we've seen how LightCrafter uses inverse rendering to recover scene properties, PBR rendering to bake illumination, and diffusion refinement for the final polish. It’s a solid framework for making videos look relit consistently over time. Lu It really pushes the boundaries of how we use generative models in a physically informed way. Meng I'm eager to see how this specific pipeline scales when we start applying it to much longer, more dynamic real-world sequences.
Tom: That’s all for today’s deep dive into LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting.
Carnegie Mellon University · University of Toronto · Bosch Research
cs.CV, cs.GR
Submitted: 2026-07-09
Updated: 2026-10-07
Importance score: 92/100
The gist: Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurate estimation of intrinsic scene properties such
Key concepts
- Scene Intrinsic Recovery
- This stage recovers essential scene information like geometry, environment maps, and surface properties (albedo/roughness) from the input video using specialized tools. Recovering geometry in a world coordinate frame is crucial for accurately calculating shadows and ensuring consistent lighting across different camera viewpoints.
- Forward PBR Rendering
- This step uses the recovered scene data to render a new video under the desired target illumination using physically-based rendering equations. Because the rendering is deterministic, it produces a full PBR video that is frame-aligned and temporally consistent by design, capturing structured lighting changes reliably.
- PBR-Conditioned Diffusion Refinement
- A video diffusion model takes the PBR render as input and refines it into a photorealistic final output. This refinement process is trained using artifact-matched pairs—synthetic renders with known ground truth—to teach the model how to correct rendering imperfections, leading to high-fidelity results.
Terminology
Summary
Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurate estimation of intrinsic scene properties such as materials, geometry, and illumination. The core contribution is LightCrafter, a hybrid pipeline that reformulates video relighting as video translation of a proxy video by translating a Physically Based Rendering (PBR) rendering under the target illumination conditions to the final target. This approach allows for baking
illumination targets into the PBR-proxy rendering, which captures most structured lighting changes and provides temporal consistency by construction, leaving a diffusion model to focus on artifact correction.
How it works
The method operates in three main stages: first, recovering scene intrinsics via inverse rendering; second, rendering a PBR video under the target illumination; and finally, refining this PBR video into a photorealistic output via a video diffusion model. This pipeline is designed to overcome the limitations of prior methods by explicitly grounding the relighting process in physical principles.
-
Scene Intrinsic Recovery: The process begins by recovering scene intrinsic properties using off-the-shelf methods, including
photometric surface properties via DiffusionRenderer [14], environment maps via DiffusionLight [23], and geometric surfaces via MegaSAM [13].
Crucially, the authors explicitly recover geometry in a world coordinate frame, enablingaccurate visibility computation for shadow rendering and consistent control over target environment maps across viewpoints.
For source illumination estimation, they initialize the environment map using DiffusionLight from the first frame and refine it via differentiable rendering to minimize photometric error against all frames of the input video. -
Forward PBR Rendering: Given the recovered intrinsics, a PBR video is rendered using a physically-based renderer [10]. The rendering equation is employed, utilizing a material model parameterized by
albedo maps a(x) and material map m(x) = (roughness and metalness).
This forward rendering produces thefull PBR video,
which is explicitly controllable because illumination enters through rendering rather than as a latent signal, frame-aligned because visibility and shading follow recovered geometry, and temporally consistent by construction since deterministic rendering applies the same target illumination uniformly across arbitrarily many frames. -
PBR-Conditioned Diffusion Refinement: The PBR rendering is then refined into a photorealistic output using a video diffusion model, initialized from CogVideoX-5B [30]. The process involves encoding both the noisy PBR rendering and the ground-truth relit video into latents, adding noise to produce a noisy latent, and training a diffusion transformer to predict the noise. To handle long videos, they employ
overlap-fused temporal tiling,
where per-window noise predictions are fused in noise-prediction space using temporal tent weights before applying a single global scheduler update to the full latent sequence, ensuringa single coherent latent trajectory over the full video.
Data Curation and Training Strategy
A novel data curation pipeline is introduced to create artifact-matched supervision. This involves two complementary sources: synthetic pairs and real-world pseudo-pairs. Synthetic pairs are constructed by rendering ground-truth videos under randomly selected target lightings, and then creating a proxy by inverse rendering the target video and re-rendering it under the target light via R to get P, exposing the model to realistic render artifacts while providing ground-truth PBR to refined video supervision.
Real-world pseudo-pairs are generated by running inverse rendering on real videos to recover intrinsics, optimizing an environment map Lˆsource such that the PBR render video matches the input frames, yielding pairs (ˆI, I) where discrepancies arise from inverse rendering errors rather than lighting change. The text emphasizes that the artifact-matched PBR renders are critical; by exposing the model to realistic rendering artifacts during training, it learns to correct in-the-wild artifacts at inference.
Scene Editing Capabilities
A key advantage of this formulation is the inherent support for downstream editing tasks. Because the refiner learns to correct imperfect PBR renders, the same trained refiner can be reused whenever an edit is applied to the scene intrinsics S or lighting Ltarget.
Edits are applied by simply rerendering an updated PBR video, such as changing material properties via UV maps, adding new lights, changing camera trajectories, and inserting new objects.
This capability is achieved because the forward renderer explicitly computes how each light source interacts with the scene.
Experimental Evaluation
LightCrafter is evaluated on synthetic and real-world benchmarks using metrics like PSNR, SSIM, LPIPS for fidelity, and T-CLIP/Warp-SSIM for temporal consistency. The experiments demonstrate that LightCrafter outperforms baselines on relighting fidelity and temporal consistency across synthetic and real-world settings.
Specifically, the method achieves state-of-the-art performance on complex scenes while enabling consistent relighting over arbitrarily long videos,
maintaining stable shadows and highlights throughout long sequences where prior methods exhibit visible degradation.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on the LightCrafter framework, along with what these improved systems will be able to do:
The proposed LightCrafter framework fundamentally enhances video relighting by replacing ambiguous generative translation or noisy inverse rendering with a structured, three-stage pipeline: Inverse Rendering (for scene structure), PBR Proxy Rendering (for physically consistent lighting control), and Diffusion Refinement (for artifact correction).
Here are the specific improvements and capabilities enabled by this system:
-
Improve the fidelity of long-form video relighting by ensuring temporal stability across arbitrary sequence lengths.
-
Enable precise, controllable manipulation of scene appearance (geometry, materials) independent of illumination changes.
-
Achieve state-of-the-art photorealism in relit videos that accurately captures intricate lighting effects like global illumination and specular reflections.
-
Allow for seamless
scene editing
via the reuse of a single trained model for various modifications (lighting, materials, object placement). -
Robustly generalize to unseen real-world video data by mitigating the train-test mismatch through an artifact-matched data curation pipeline.
Specific improvements and enhanced capabilities:
-
A photorealistic video relighting system capable of maintaining perfect temporal consistency over arbitrarily long videos (e.g., hours of footage) without the drift or chunking artifacts seen in current end-to-end diffusion models.
-
A system that allows users to change material properties (e.g., changing a surface from matte plastic to polished metal) or insert objects into a scene, and have the relighting model automatically harmonize the new geometry/materials with the target illumination, all without requiring retraining of the core diffusion model.
-
A system that accurately models and renders complex lighting scenarios, including subtle global illumination (indirect light bounces), sharp shadows, highlights, and reflections on metallic or rough surfaces. This moves beyond simple environment map mapping to physically grounded light transport simulation.
-
A system that can perform
scene editing
by inserting virtual light sources (e.g., a specific spotlight or an area light) into the scene and having the diffusion refiner produce a photorealistic result that correctly interacts with this new light source, while maintaining fidelity to the original scene structure. -
A robust inference pipeline capable of processing videos of any length by using overlap-fused temporal tiling, ensuring that lighting and shadows remain coherent across scene boundaries, even when dealing with sequences much longer than the model's training window (e.g., 200+ frames).
-
A system that achieves superior generalization to real-world
in-the-wild
footage by learning to correct artifacts specific to the PBR rendering pipeline during training, making it less prone to hallucinating incorrect structures or failing on complex geometries like foliage and thin objects compared to models trained only on synthetic data.
Sources
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- Objaverse: A Universe of Annotated 3D Objects
- MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
- LuxRemix: Lighting Decomposition and Remixing for Indoor Scenes
- Decoupled Weight Decay Regularization
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models
- DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models