GAN-Diff: Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets

arXiv:2608.22272 · cs.CV, cs.AI, cs.LG · Submitted 2026-08-23 · Read on arXiv

cs.CV, cs.AI, cs.LG

Submitted: 2026-08-23

Updated: 2026-08-27

Comments: 7 pages, 10 figures, 4 tables

License: http://creativecommons.org/publicdomain/zero/1.0/

The gist: Generative adversarial networks (GANs) can provide efficient image generation, while diffusion models offer high-quality image restoration but require iterative sampling.

Terminology

Abstract

Generative adversarial networks (GANs) can provide efficient image generation, while diffusion models offer high-quality image restoration but require iterative sampling. This paper presents a hybrid GAN-guided diffusion framework that uses a pretrained Wasserstein GAN with gradient penalty (WGAN-GP) as a feature prior for conditional diffusion-based image restoration. Intermediate features from the frozen WGAN-GP generator are incorporated into a diffusion U-Net through cross-attention and remain fixed during the DDIM sampling process. The framework is evaluated on two restoration tasks, Gaussian denoising and 2Xsuper-resolution, using CelebA face images. During development, several sources of instability were identified and addressed, including adversarial learning-rate imbalance, inappropriate diffusion initialization, excessive corruption, and insufficient parameter averaging. The resulting framework consistently improves the quality of both degraded and low-resolution images. In particular, it improves denoising performance by 4.40 dB in PSNR and super-resolution performance by 3.70 dB over their respective input baselines. These results demonstrate the potential of a frozen GAN feature prior to guide diffusion models toward stable and effective image restoration.

Sources

Related papers