pi-SUB: A Physics-Informed Synthetic Underwater Benchmark Dataset for Underwater Image Enhancement

arXiv:2608.10589 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Namritha Lasyapriya Maddali, Rajini Makam, Suresh Sundaram, Narasimhan Sundararajan

Indian Institute of Science · PES University · Nanyang Technological University

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 13 pages

Code: https://github.com/airl-iisc/pi-SUB

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper presents π-SUB, a physics-informed framework for generating synthetic underwater benchmark datasets that bridges the synthetic-to-real gap for Underwater Image Enhancement (UIE).

Terminology

Summary

This paper presents π-SUB, a physics-informed framework for generating synthetic underwater benchmark datasets that bridges the synthetic-to-real gap for Underwater Image Enhancement (UIE). The proposed framework extends the classical underwater image formation model by incorporating depth-dependent downwelling irradiance, biologically resolved absorption, and environmental scattering across all ten Jerlov water types, together with independently controllable residual phenomena. Using this framework, the π-SUB dataset consists of paired synthetic underwater–reference images spanning shallow-to-deep and coastal-to-oceanic environments. Extensive simulation studies have been carried out to evaluate π-SUB along two criteria namely hyper-realism and generalizability. For hyper-realism, π-SUB attains a global Fréchet Inception Distance (FID) that is 46% lower than Syrea. For generalizability, four state-of-the-art UIE architectures (FUnIE-GAN, Pix2Pix, PUIE-Net, and Phaseformer) are used for comparative evaluation of π-SUB. These models were independently trained on six datasets including one real and five synthetic datasets and tested on six real-world benchmarks datasets. Across four UIE architectures and six real benchmark datasets, π-SUB improves UIQM by 4.18% over PHISWID (next best) and 9.46% over Syrea (next best), while reducing NIQE by 48.78% and 23.98%, respectively. These results establish π-SUB as a hyper-realistic and generalizable benchmark for developing the next generation of underwater image enhancement methods.

The main contributions of the paper are: (i) A physics-informed framework, π-SUB, has been developed to generate paired underwater benchmark datasets from either a set of clean real images or Unreal Engine rendered images by modeling the following, viz the water type, the camera depth, visibility, chlorophyll concentration, and finally optical attenuation. (ii) A modified Jaffe–McGlamery image formation model is developed that incorporates the depth-dependent downwelling irradiance, Jerlov IOPs, separate direct-transmission and backscatter attenuation coefficients, and finally biologically driven optical effects including chlorophyll absorption, CDOM, and fluorescence. (iii) Extensive performance evaluation of π-SUB has shown that it produces a hyper-realistic and highly generalizable training data, achieving a 46% lower FID than the next best synthetic benchmark dataset Syrea and consistently improving UIE performance across all other state-of-the-art architectures and real-world benchmark datasets.

The framework accepts three user-defined inputs: (i) a reference image source, selected from either a simulated or real image pool; (ii) the physical water configuration, specified by the Jerlov water type, camera depth, and the corresponding inherent and apparent optical properties (IOPs/AOPs); and (iii) an Augmented Realism configuration, which determines whether suspended particulate matter, volumetric haze, biological effects, or their combination are applied to further degrade the images. The framework then proceeds through three sequential stages. Stage I estimates a dense scene-range map and converts it to a metric underwater range using the water-type-dependent maximum visibility distance zmax. Stage II evaluates the proposed modified underwater image formation model to synthesize the deterministic underwater observation Ib. Finally, Stage III uses the selected residual optical phenomena to generate the phenomenon-specific and fully augmented subsets that together constitute the π-SUB dataset.

The modified underwater image formation model is given by: Ib(λ) = J Ed(λ)e(−βd(λ)zs) + B∞(λ)(1 − e(−βb(λ)zs)), where Ib is referred as Baseline image, Ed = E0 e(−Kd d), βd = a + δb, βb = a + b, B∞ = bb Ed/(a + b), and a = aw + achl + acdom. Unlike the conventional Jaffe–McGlamery model, the proposed formulation explicitly separates direct and backscatter attenuation, incorporates depth-dependent illumination through Ed, and resolves absorption into measured biological constituents while remaining deterministic and analytically invertible.

The π-SUB comprises a deterministic MUIF subset (Ib) together with phenomenon-specific variants: MUIF+Haze (Ih), MUIF+BE (Ibp), and MUIF+SP (Isp), as well as a fully augmented subset, MUIF+Haze+BE+SP (Ic), which combines all three residual phenomena. The released benchmark is generated from N = 117 clean reference images drawn from the simulated and real pools. Each reference is synthesized under all ten Jerlov water types for the Ib, Isp, and Ibp subsets, under ten water types at three severity settings for Ih and Ic, and under thirty randomized effect–depth–water-type combinations for the held-out test split, giving 117 × 120 = 14,040 paired synthetic underwater–reference images in total.

For hyper-realism validation, π-SUB achieves the lowest global FID at 95, compared to 176 for Syrea, 192 for SUIEB, 227 for PHISWID, and 230 for SUID. Only 2.00% of the 14,040 π-SUB images are flagged as out-of-distribution, against 4.02% for PHISWID, 12.23% for Syrea, 19.26% for SUIEB, and 27.33% for SUID. π-SUB is the only synthetic dataset represented in all seven perceptual clusters obtained from K-means clustering of real underwater imagery.

For generalizability, twenty-four models (four architectures × six training datasets) are evaluated on six real benchmarks. Models trained on π-SUB rank first or second in UIQM on UIEB, OceanDark, and SQUID in all twenty-four architecture–benchmark combinations shown, the best win rate of any candidate training set on these three benchmarks. Averaged across four UIE architectures and six real benchmark datasets, models trained on π-SUB improve UIQM by 9.46% over Syrea and 4.18% over PHISWID, while reducing NIQE by 23.98% and 48.78%, respectively. In a downstream feature-matching task on consecutive SQUID Katzaa frames, the raw pair yields only 23 verified matches, PhaseFormer trained on PHISWID raises this to 85, while training on π-SUB yields 157: a 6.8× improvement over the raw frames and 1.8× over PHISWID.

Improvements for AI systems

Improvements to AI systems:

  1. Physics-constrained generative training: Train UIE models with a loss term that penalizes deviation from the modified Jaffe–McGlamery forward model (depth-dependent downwelling irradiance, separate direct/backscatter attenuation, biologically resolved absorption). This forces the network to learn physically consistent mappings rather than purely statistical ones, improving robustness to unseen water types and depths.

  2. Jerlov-aware domain adaptation: Implement a conditioning mechanism (e.g., FiLM layers or adaptive instance normalization) that takes Jerlov water type, camera depth, and chlorophyll concentration as input. This allows a single model to dynamically adjust its enhancement behavior across all ten water types, replacing the need for separate models per environment.

  3. Residual phenomenon decomposition: Train a multi-head architecture that separately predicts and removes each degradation component (haze, biological effects, suspended particles) before fusing them into the final enhanced image. This mirrors π-SUB’s subset structure (Ib, Ih, Ibp, Isp, Ic) and enables targeted correction, improving interpretability and performance on partially degraded real images.

  4. Uncertainty-aware enhancement via depth estimation: Use the dense scene-range map generation from Stage I as a pretraining task for a depth-estimation head. The enhanced image can then be produced by applying the inverse of the forward model using the estimated depth, with the network only learning residual corrections—reducing the search space and improving generalization to unseen depths.

  5. Synthetic-to-real feature alignment: During training, add a perceptual loss computed against features from a backbone pretrained on real underwater images (e.g., using the K-means clusters identified in π-SUB). This aligns the synthetic training distribution with real-world perceptual clusters, reducing the domain gap and improving downstream feature matching (e.g., 6.8× improvement in feature matches on SQUID frames).

What the improved AI system can do:

  • Enhance underwater images from any of the ten Jerlov water types, at arbitrary depths, with a single unified model—without retraining or fine-tuning per environment.

  • Produce physically consistent enhancements that remain valid under varying chlorophyll, CDOM, and particulate concentrations, even in conditions not seen during training.

  • Decompose the degradation into interpretable components (haze, biological, particles), allowing users to selectively remove only the artifacts present in a given image.

  • Achieve state-of-the-art performance on real-world benchmarks (UIQM, NIQE) while being more robust to distribution shift, as evidenced by lower out-of-distribution rates and higher feature-matching success in downstream tasks like visual odometry or 3D reconstruction.

  • Provide confidence estimates for the enhancement by leveraging depth uncertainty, enabling safer deployment in autonomous underwater vehicles or remote sensing applications where incorrect enhancement could lead to navigation errors.

Abstract

This paper presents pi-SUB, a physics-informed framework for generating synthetic underwater benchmark datasets that bridges the synthetic-to-real gap for Underwater Image Enhancement (UIE). The proposed framework extends the classical underwater image formation model by incorporating depth-dependent downwelling irradiance, biologically resolved absorption, and environmental scattering across all ten Jerlov water types, together with independently controllable residual phenomena. Using this framework, the pi-SUB dataset consists of paired synthetic underwater-reference images spanning shallow-to-deep and coastal-to-oceanic environments. Extensive simulation studies have been carried out to evaluate pi-SUB along two criteria namely hyper-realism and generalizability. For hyper-realism, pi-SUB attains a global Frechet Inception Distance (FID) that is 46% lower than Syrea. For generalizability, four state-of-the-art UIE architectures (FUnIE-GAN, Pix2Pix, PUIE-Net, and Phaseformer) are used for comparative evaluation of pi-SUB. These models were independently trained on six datasets including one real and five synthetic datasets and tested on six real-world benchmarks datasets. Across four UIE architectures and six real benchmark datasets, pi-SUB improves UIQM by 4.18% over PHISWID (next best) and 9.46% over Syrea (next best), while reducing NIQE by 48.78% and 23.98%, respectively. These results establish pi-SUB as a hyper-realistic and generalizable benchmark for developing the next generation of underwater image enhancement methods. The code and dataset are available at https://github.com/airl-iisc/pi-SUB

Sources

Related papers