From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
Zepeng Wang, Jiagao Hu, Fuhao Li, Yuxuan Chen, Fei Wang, Daiguo Zhou
Xiaomi Inc.
cs.CV, cs.AI, eess.IV
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: Project page: https://codingwzp.github.io/VideoDereflection_S2R
Code: https://github.com/huggingface/controlnet_aux
Project page: https://codingwzp.github.io/VideoDereflection_S2R
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 55/100
The gist: This paper presents a closed-loop framework for video reflection removal, unifying physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation.
Terminology
Summary
This paper presents a closed-loop framework for video reflection removal, unifying physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation. The authors identify that video reflection removal is underexplored due to a lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks.
The paper proposes three main contributions:
-
S2R-Synthesis: A paired video reflection synthesis pipeline that generates realistic and controllable reflected videos. It decouples reflection structure from appearance by using a video diffusion renderer (built on Wan2.1) that renders photorealistic reflections from a clean transmission video and a structured reflection condition. A Physics-Grounded Augmentation (PGA) module transforms the reflection condition according to glass optical parameters, controlling effects like surface roughness, glass reflectance, and thickness.
-
S2R-Removal: The first diffusion-based video reflection removal model. It adapts a pretrained video diffusion prior (Wan2.1) through a two-stage training process. Stage I performs reflection-aware latent adaptation with reflection-intensity supervision to localize and remove reflections. Stage II refines the model with pixel-geometric losses (reconstruction, structural, and depth consistency), training it to recover the clean transmission in a single denoising step. This one-step design makes it faster than non-diffusion baselines.
-
S2R-Bench: The first benchmark for video reflection removal, containing two subsets: S2R-Ref for full-reference evaluation with paired videos and S2R-Real for real-world human perceptual assessment.
The paper states: We present a closed-loop framework that unifies physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation.
The experiments demonstrate state-of-the-art performance and faster inference than even non-diffusion baselines
on S2R-Bench and multiple public image benchmarks. The ablation study confirms the effectiveness of the S2R-Synthesis pipeline, showing that Full PGA reaches 34.13 dB PSNR and 0.9704 SSIM, a gain of +1.78 dB and +0.0061 SSIM over the planar-only setting.
The removal model achieves 28.84 dB PSNR and 0.8594 SSIM
on S2R-Ref with its full objective.
Improvements for AI systems
Improvements to AI systems:
-
Physics-grounded synthetic data generation for video tasks – The S2R-Synthesis pipeline can be integrated into any video restoration or generation model (e.g., dehazing, deblurring, low-light enhancement) to create large-scale, controllable paired training data. The Physics-Grounded Augmentation (PGA) module enables precise manipulation of optical parameters (roughness, reflectance, thickness), allowing AI systems to learn robust, physically consistent transformations rather than relying on unrealistic synthetic artifacts.
-
Single-step diffusion inference for real-time video processing – The two-stage training (reflection-aware latent adaptation + pixel-geometric refinement) that yields a one-step denoising diffusion model can be applied to other video-to-video translation tasks (e.g., super-resolution, denoising, frame interpolation). This reduces inference latency from multi-step diffusion (typically 20–50 steps) to a single step, making diffusion-based systems viable for real-time or near-real-time applications without sacrificing quality.
-
Temporally coherent removal/restoration in video diffusion priors – The reflection-intensity supervision in Stage I and depth-consistency loss in Stage II can be reused to enforce temporal stability and geometric fidelity in any video diffusion model. This improves the AI’s ability to maintain consistent object boundaries, lighting, and motion across frames, addressing a common failure mode of per-frame processing.
-
Closed-loop evaluation with human perceptual benchmarks – The dual-benchmark design (S2R-Ref for full-reference metrics, S2R-Real for human assessment) provides a template for evaluating any video enhancement AI. By incorporating both quantitative and perceptual criteria, AI systems can be tuned to match human visual preferences, not just PSNR/SSIM, leading to more practically useful outputs.
-
Reflection-conditioned rendering for controllable content generation – The video diffusion renderer that decouples reflection structure from appearance can be adapted for other controllable generation tasks (e.g., inserting watermarks, lens flares, or dynamic lighting effects into videos). This enables AI systems to generate physically plausible optical effects with user-specified parameters, useful for film production, AR/VR, and synthetic data creation.
What the improved AI system can do:
-
Real-time video reflection removal on live camera feeds or video calls, with quality exceeding prior non-diffusion methods and latency low enough for interactive use.
-
Generate physically accurate synthetic training data for any optical degradation (reflections, glare, haze) with user-controlled glass properties, enabling rapid domain adaptation to new camera hardware or environments.
-
Perform temporally stable video restoration (e.g., deblurring, denoising) in a single diffusion step, preserving object identity and motion coherence across frames.
-
Create controllable visual effects in videos—such as adding realistic reflections to a clean scene or adjusting glass thickness/roughness—for content creation and simulation.
-
Evaluate and optimize video AI models using both automated metrics and human perceptual scores, ensuring the system’s outputs are not only mathematically accurate but also visually pleasing to end users.
Abstract
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection removal remains largely underexplored due to the lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks. We present a closed-loop framework that unifies physics-grounded reflection simulation, diffusion-based video dereflection, and benchmark evaluation. Our S2R-Synthesis pipeline generates paired reflected and reflection-free videos by performing physics-grounded augmentation in the structure space and rendering realistic reflected videos with a trained video diffusion renderer; the augmentation models key glass-related effects including roughness-induced blur, thickness-induced ghosting, and reflectance variation. Based on the synthesized data, we introduce S2R-Removal, the first diffusion-based video reflection removal model, which adapts a pretrained video diffusion prior through reflection-aware latent adaptation and one-step pixel-geometric refinement, recovering the clean transmission in a single denoising step. We further build S2R-Bench, the first benchmark for video reflection removal, supporting both full-reference evaluation and real-world human perceptual assessment. Experiments on S2R-Bench and multiple public image benchmarks demonstrate state-of-the-art performance and faster inference than even non-diffusion baselines, and validate the effectiveness of S2R-Synthesis. Project page: https://codingwzp.github.io/VideoDereflection S2R.
Sources
- Rectifying Latent Space for Generative Single-Image Reflection Removal
- Wan: Open and Advanced Large-Scale Video Generative Models
- Reflection Removal through Efficient Adaptation of Diffusion Transformers
- PromptRR: Diffusion Models as Prompt Generators for Single Image Reflection Removal
- FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal
- Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
- Qwen3-VL Technical Report
- UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
- Kling-Omni Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models