From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection

arXiv:2608.11562 · cs.CV, cs.AI, eess.IV · Submitted 2026-08-12 · Read on arXiv

Zepeng Wang, Jiagao Hu, Fuhao Li, Yuxuan Chen, Fei Wang, Daiguo Zhou

Xiaomi Inc.

cs.CV, cs.AI, eess.IV

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Project page: https://codingwzp.github.io/VideoDereflection_S2R

Code: https://github.com/huggingface/controlnet_aux

Project page: https://codingwzp.github.io/VideoDereflection_S2R

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 55/100

The gist: This paper presents a closed-loop framework for video reflection removal, unifying physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation.

Terminology

Summary

This paper presents a closed-loop framework for video reflection removal, unifying physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation. The authors identify that video reflection removal is underexplored due to a lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks.

The paper proposes three main contributions:

  1. S2R-Synthesis: A paired video reflection synthesis pipeline that generates realistic and controllable reflected videos. It decouples reflection structure from appearance by using a video diffusion renderer (built on Wan2.1) that renders photorealistic reflections from a clean transmission video and a structured reflection condition. A Physics-Grounded Augmentation (PGA) module transforms the reflection condition according to glass optical parameters, controlling effects like surface roughness, glass reflectance, and thickness.

  2. S2R-Removal: The first diffusion-based video reflection removal model. It adapts a pretrained video diffusion prior (Wan2.1) through a two-stage training process. Stage I performs reflection-aware latent adaptation with reflection-intensity supervision to localize and remove reflections. Stage II refines the model with pixel-geometric losses (reconstruction, structural, and depth consistency), training it to recover the clean transmission in a single denoising step. This one-step design makes it faster than non-diffusion baselines.

  3. S2R-Bench: The first benchmark for video reflection removal, containing two subsets: S2R-Ref for full-reference evaluation with paired videos and S2R-Real for real-world human perceptual assessment.

The paper states: We present a closed-loop framework that unifies physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation. The experiments demonstrate state-of-the-art performance and faster inference than even non-diffusion baselines on S2R-Bench and multiple public image benchmarks. The ablation study confirms the effectiveness of the S2R-Synthesis pipeline, showing that Full PGA reaches 34.13 dB PSNR and 0.9704 SSIM, a gain of +1.78 dB and +0.0061 SSIM over the planar-only setting. The removal model achieves 28.84 dB PSNR and 0.8594 SSIM on S2R-Ref with its full objective.

Improvements for AI systems

Improvements to AI systems:

  1. Physics-grounded synthetic data generation for video tasks – The S2R-Synthesis pipeline can be integrated into any video restoration or generation model (e.g., dehazing, deblurring, low-light enhancement) to create large-scale, controllable paired training data. The Physics-Grounded Augmentation (PGA) module enables precise manipulation of optical parameters (roughness, reflectance, thickness), allowing AI systems to learn robust, physically consistent transformations rather than relying on unrealistic synthetic artifacts.

  2. Single-step diffusion inference for real-time video processing – The two-stage training (reflection-aware latent adaptation + pixel-geometric refinement) that yields a one-step denoising diffusion model can be applied to other video-to-video translation tasks (e.g., super-resolution, denoising, frame interpolation). This reduces inference latency from multi-step diffusion (typically 20–50 steps) to a single step, making diffusion-based systems viable for real-time or near-real-time applications without sacrificing quality.

  3. Temporally coherent removal/restoration in video diffusion priors – The reflection-intensity supervision in Stage I and depth-consistency loss in Stage II can be reused to enforce temporal stability and geometric fidelity in any video diffusion model. This improves the AI’s ability to maintain consistent object boundaries, lighting, and motion across frames, addressing a common failure mode of per-frame processing.

  4. Closed-loop evaluation with human perceptual benchmarks – The dual-benchmark design (S2R-Ref for full-reference metrics, S2R-Real for human assessment) provides a template for evaluating any video enhancement AI. By incorporating both quantitative and perceptual criteria, AI systems can be tuned to match human visual preferences, not just PSNR/SSIM, leading to more practically useful outputs.

  5. Reflection-conditioned rendering for controllable content generation – The video diffusion renderer that decouples reflection structure from appearance can be adapted for other controllable generation tasks (e.g., inserting watermarks, lens flares, or dynamic lighting effects into videos). This enables AI systems to generate physically plausible optical effects with user-specified parameters, useful for film production, AR/VR, and synthetic data creation.

What the improved AI system can do:

  • Real-time video reflection removal on live camera feeds or video calls, with quality exceeding prior non-diffusion methods and latency low enough for interactive use.

  • Generate physically accurate synthetic training data for any optical degradation (reflections, glare, haze) with user-controlled glass properties, enabling rapid domain adaptation to new camera hardware or environments.

  • Perform temporally stable video restoration (e.g., deblurring, denoising) in a single diffusion step, preserving object identity and motion coherence across frames.

  • Create controllable visual effects in videos—such as adding realistic reflections to a clean scene or adjusting glass thickness/roughness—for content creation and simulation.

  • Evaluate and optimize video AI models using both automated metrics and human perceptual scores, ensuring the system’s outputs are not only mathematically accurate but also visually pleasing to end users.

Abstract

Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection removal remains largely underexplored due to the lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks. We present a closed-loop framework that unifies physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation. Our S2R-Synthesis pipeline generates paired reflected and reflection-free videos by performing physics-grounded augmentation in the structure space and rendering realistic reflected videos with a trained video diffusion renderer; the augmentation models key glass-related effects including roughness-induced blur, thickness-induced ghosting, and reflectance variation. Based on the synthesized data, we introduce S2R-Removal, the first diffusion-based video reflection removal model, which adapts a pretrained video diffusion prior through reflection-aware latent adaptation and one-step pixel-geometric refinement, recovering the clean transmission in a single denoising step. We further build S2R-Bench, the first benchmark for video reflection removal, supporting both full-reference evaluation and real-world human perceptual assessment. Experiments on S2R-Bench and multiple public image benchmarks demonstrate state-of-the-art performance and faster inference than even non-diffusion baselines, and validate the effectiveness of S2R-Synthesis. Project page: https://codingwzp.github.io/VideoDereflection S2R.

Sources

Related papers