Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration
Sangwoo Jo, Donggeun Ko, Jayeon Kang, Youngsang Kwak, Jaehwa Kwak, Sungjoon Choi
Korea University · Aim Future
cs.CV, cs.AI, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: Accepted to ECCV 2026. Code is available at https://github.com/aiimaginglab/PCFlow
Code: https://github.com/aiimaginglab/PCFlow
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: Image restoration is fundamentally constrained by the tradeoff between distortion and perception: "minimizing pixel-wise error yields over-smoothed results, whereas optimizing for perceptual realism
Terminology
Summary
Image restoration is fundamentally constrained by the tradeoff between distortion and perception: minimizing pixel-wise error yields over-smoothed results, whereas optimizing for perceptual realism often introduces structural deviations.
Recent approaches attempt to balance this tradeoff via posterior sampling or multi-stage generative pipelines, yet remain computationally expensive and architecturally complex.
The authors propose PCFlow (Perceptually Consistent Flow Matching), a unified framework that directly parameterizes a continuous transport from degraded observations to clean targets, jointly optimizing distortion and perceptual quality.
Instead of performing stochastic posterior sampling or decomposing restoration into multiple stages, PCFlow tackles the multi-objective optimization directly within a continuous latent space.
1. Latent Consistency Flow Matching (LCFM): PCFlow parameterizes a direct vector field from the degraded input to the clean target
using a latent consistency flow matching objective that enforces geometric consistency along the trajectory, effectively straightening the integration path and enabling image restoration in as few as three inference steps.
The LCFM objective penalizes discrepancies in both the predicted trajectory endpoints and the velocity fields.
2. Latent Consistency Perceptual Loss (LCPL): Recognizing that learning a few-step transport with an L2 based objective inherently risks regression to the posterior mean,
the authors introduce LCPL which encourages the trajectory endpoints to align with perceptually sharp, high-density data manifolds.
This includes both external perceptual supervision (E-LatentLPIPS) and internal perceptual supervision using the model's own decoder features.
3. Conflict-Free Gradient Alignment: The authors analyze that simply combining these objectives introduces gradient conflicts, particularly at early transport timesteps (low SNR regimes).
They adopt "a conflict-free gradient projection strategy along with an SNR-adaptive schedule, using the perceptual objective as a steering signal while removing structural gradient components that conflict with perceptual optimization." The update rule preserves the perceptual gradient and removes only the conflicting component of the structural gradient.
4. Lightweight Architecture: PCFlow employs a lightweight convolution-only model architecture without attention modules, significantly reducing model size and computational cost.
The model uses a pretrained Tiny AutoEncoder with a convolution-only U-Net backbone.
PCFlow uses a two-stage training procedure: "Initially, the model focuses solely on reconstruction quality by setting λLCPL = 0 for the first 250 training epochs. In the second stage, we enable the perceptual objective and continue training for an additional 250 epochs with the proposed SNR-adaptive weighting." The perceptual weight λLCPL(t) follows a linear warmup schedule that increases with timestep t.
Blind Face Restoration (BFR): PCFlow achieves state-of-the-art perceptual quality, obtaining the best FID and NIQE on CelebA-Test and the best FID on CelebAdult, while requiring only 32M parameters and five sampling steps (K = 5).
Compared to ELIR, PCFlow uses fewer parameters (32M vs. 37.5M) and achieves 1.29× higher inference speed (42.62 vs. 33.11 FPS), while substantially improving FID (35.89 vs. 44.64) on CelebA-Test.
Compared to PMRF (183M parameters, 0.57 FPS), PCFlow outperforms in perceptual quality metrics while delivering over 75× higher throughput.
Other Restoration Tasks: PCFlow consistently improves over ELIR in FID across all four image restoration tasks
(super-resolution, denoising, inpainting, colorization), using only 21M parameters compared to ELIR (27M) and PMRF (176M).
-
Preheating period is essential:
Jointly training both distortion and perception objectives from the beginning yields inferior results compared to PCFlow.
-
Conditional flow formulation improves FID substantially (from 56.27 to 47.21 on colorization).
-
Encoder fine-tuning consistently improves both reconstruction and perceptual quality.
-
Linear warmup λ-scheduling achieves the best FID across all tasks.
-
Internal perceptual network with conflict-free gradient alignment yields the best perceptual quality.
-
Projecting the structural gradient (rather than perceptual gradient) yields better FID.
The authors state: "PCFlow establishes a more favorable distortion–perception trade-off frontier compared to previous baselines across various image restoration tasks, while requiring only a few inference steps and maintaining an efficient model architecture. The method demonstrates that
directly learning the conditional transport from degraded observations, combined with perceptual supervision as a steering signal, is sufficient to achieve superior distortion-perception trade-offs under a few-step regime."
Improvements for AI systems
Improvements to AI systems:
-
Few-step generative restoration with perceptual steering: Implement a latent consistency flow matching objective that directly parameterizes a continuous transport from degraded to clean images, enabling restoration in as few as 3–5 inference steps. Add a latent consistency perceptual loss (LCPL) that aligns trajectory endpoints with perceptually sharp manifolds, using both external (LatentLPIPS) and internal (decoder feature) supervision. The improved system can perform blind face restoration, super-resolution, denoising, inpainting, and colorization with state-of-the-art FID/NIQE while running at 42+ FPS on a single GPU—over 75× faster than prior generative baselines like PMRF.
-
Conflict-free multi-objective optimization with SNR-adaptive scheduling: Replace naive loss weighting with a gradient projection strategy that preserves the perceptual gradient and removes only the conflicting component of the structural gradient, combined with a linear warmup schedule for the perceptual weight that increases with transport timestep (low SNR = lower weight). The improved system avoids the distortion–perception tradeoff failure modes (over-smoothing or structural hallucination) and achieves a better Pareto frontier, as validated by ablation showing joint training from scratch degrades FID by 9 points.
-
Lightweight convolution-only architecture with pretrained autoencoder: Use a convolution-only U-Net backbone (no attention) with a pretrained Tiny AutoEncoder, reducing parameters to 21–32M (vs. 176–183M for PMRF). The improved system can be deployed on edge devices or real-time applications (e.g., video conferencing, mobile photo editing) without sacrificing perceptual quality, while maintaining a 1.29× speed advantage over ELIR and 75× over PMRF.
-
Two-stage training with preheating: First train 250 epochs on reconstruction only (λLCPL=0), then enable perceptual loss for another 250 epochs with SNR-adaptive weighting. The improved system converges more stably and avoids early gradient conflicts, leading to consistently better FID across all four restoration tasks compared to joint training from scratch.
-
Conditional flow formulation with encoder fine-tuning: Model the transport as conditional on the degraded input (not unconditional), and fine-tune the encoder during training. The improved system gains substantial FID improvements (e.g., from 56.27 to 47.21 on colorization) and better reconstruction fidelity, enabling robust performance on diverse degradation types without task-specific retraining.
Sources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models