PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation
Elad Yoshaia, Natan T. Shaked
School of Biomedical Engineering, Faculty of Engineering, Tel Aviv University
cs.CV, cs.AI
Submitted: 2026-08-06
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 69/100
The gist: The paper addresses unpaired image-to-image translation, which "maps an image from a source domain to a target domain while preserving its semantic content" but is "fundamentally under-constrained"
Terminology
Summary
The paper addresses unpaired image-to-image translation, which maps an image from a source domain to a target domain while preserving its semantic content
but is fundamentally under-constrained
without paired supervision. The central difficulty is control: deciding, per image and per region, which features should change and which must be kept.
The authors criticize existing diffusion-based translators: Many diffusion-based unpaired translators control preservation through a single global noise or guidance value applied across the image, which cannot separate content to keep from appearance to change.
As they state, "SDEdit exposes a single global knob, the noise level, which governs both how much structure is destroyed and how much style may change... the trade-off is applied uniformly to every pixel. Real translations are heterogeneous. Staining must change while nuclear outlines must not, so the optimal control is a field rather than a scalar."
PRISM is described as a GAN-free flow-matching framework that replaces this global control with a learned per-feature gate.
The authors recast unpaired translation as selective per-feature transport to the target distribution,
where each feature should change in proportion to how far its own feature statistics sit from the target distribution, neither too much (which hallucinates) nor too little (which under-translates).
-
A distribution-informed per-feature preservation gate (DDτ):
A spatial prior is set from the feature's standardized distance to the target distribution (with a per-image robust normalization) and supplied as a soft target for a learned channel-spatial gate under joint training.
The gate isoverridable at inference from a text prompt or a domain detector without retraining.
-
The same gate controls both initialization and transport timing: "High-τ features are initialized from the real source latent and receive near-zero transport, so preserved content stays at its source value, and the same gate governs when each feature is allowed to move during integration."
-
Task-matched initialization:
A single design switch at t=0, AdaIN(source content, target style) for structure-preserving tasks and a partially-anchored isotropic corruption for structure-changing ones, adapts the framework to both regimes.
Stage 1 learns a generative flow that maps Gaussian noise to each domain, which Stage 2 then reuses frozen.
Using the linear-interpolant flow matching of [20], a domain-conditional network vθ(zt, t, d) is trained with the flow-matching loss: LFM = E[∥vθ(zt, t, d) − (z − ε)∥2]. After training, vθ is frozen, meaning MB is now fixed, so Stage 2 optimizes against a stationary target rather than a moving one.
PRISM uses "a per-feature gate field τ ∈ τmin, τmax... one scalar per latent feature-location. At each location it
quantifies how strongly the feature should be preserved (τ→1) versus translated (τ→0). The idealized transport target is formalized as
T⋆(xA) = arg min over x̂: E(x̂)∈MB of Σp τ(p) E(x̂)(p) − E(xA)(p)2, where
High τ is expensive and preserved, low τ cheap and freed."
A small (3M-parameter) U-Net Tψ reads the source latent and emits: τ = τmin + (τmax − τmin) σ(Tψ(zA)),
with τmin≈0.05, τmax=1. We zero-initialize Tψ, so training starts from a neutral gate τinit everywhere.
The gate enters in two places:
1. Initialization: z(0) = τ ⊙ zA + (1 − τ) ⊙ εα
— a per-feature generalization of SDEdit. Where SDEdit renoises the whole image by one global amount, Equation (4) renoises each feature by its own amount 1 − τ.
2. Integration (wake-up timing): vgated = σ((tk − τ)/T) ⊙ v, z(k+1) = z(k) + ∆t·vgated
where a high-τ feature remains inactive through most of the trajectory, while a low-τ feature becomes active earlier and translates sooner.
The sharpness T=0.15 trades gradient flow against gate crispness.
The corruption is task-matched. For structure-preserving tasks (Type 1), the content-anchored corruption is: εstruct = IN(zA) ⊙ σ̃ + µ̃,
which is exactly Adaptive Instance Normalization (AdaIN).
For structure-changing tasks (Type 2), isotropic noise is used, and the two are blended: εα = α εstruct + (1 − α) ε.
For Type-1 tasks α=1 (with a one-sided contrast floor σ̃ ≥ c σc, c=0.85
), for Type-2 tasks α=0.5, and virtual staining uses α=0.3. The paper notes α thus exposes a controllable trade-off between distribution fidelity and input faithfulness.
The wake-up gate enables a built-in efficiency opportunity
: NFE(p) = (1 − τ(p))/∆t,
so preserved (high-τ) features cost few function evaluations and translated features receive the full budget.
The authors caution these are idealized per-region evaluation counts, not wall-clock savings
because the dense backbone still evaluates every spatial position at each step.
A 17M-parameter DiT Cϕ outputs an additive velocity vC = Cϕ(z(k), tk, ϕfeat(xA), dir, cdiag), zero-initialized and norm-constrained,
with the clipping vC ← vC · min(1, β∥vF∥/(∥vC∥ + εnum)), β = 0.5.
This can adjust the frozen vector field to counter off-manifold drift while its magnitude stays bounded relative to the base velocity, without overwriting the base flow, which supports stability without an adversary.
For non-saturating, per-region realism,
PRISM caches for each source, the statistics of its k nearest real targets in frozen DINOv2 space
and minimizes a per-sample Mahalanobis distance.
Each source has its own target statistics,
translating this specific cat toward the target region that looks most like it.
The gate's spatial pattern is supervised by: d(p) = ∥(ϕ(xA)(p) − µB)/σB∥2, τprior(p) = 1 − normq(d(p)),
with per-image q-quantile normalization where q is the 0.95-quantile. The loss is LDDτ = ∥τ − sg(τprior)∥2.
The paper explains: patches that are already close to the target need little transport, while distant patches require more.
For medical imaging, Stage 1 is conditioned on organ/diagnosis/grade metadata
via cdiag injected into vθ, and Stage 2 adds a diagnostic-consistency loss using a frozen diagnosis classifier hψ
with Ldiag = CE(hψ(x̂B), arg max hψ(xA)).
L = λmmd Lmmd + λDDτ LDDτ + λtv Ltv + λspread Lspread + λstruct Lstruct + λreg Lreg,
decomposed into a realism term that pulls outputs onto the target set, a gate term that supervises the measured per-feature gate, and for structure-preserving tasks only, a preservation anchor.
Lmmd is a multi-layer MMD; Ltv and Lspread keep τ spatially coherent and well-conditioned; Lstruct is a τ-gated structure loss with DINOv2 patch, L1, and Sobel-gradient distances; supporting terms include Ldom, Llocal match, and Ltype. No adversarial or PatchGAN loss is used.
To remove seed dependence in Type-2 tasks, A small generator Gξ... emits, per image, a single deterministic ε⋆, so the translation becomes a reproducible one-to-one map,
written as ε⋆ = εstruct + δ Gξ(zA)
with a distribution-validity penalty Lndist = ∥µc(ε⋆) − µB∥2 + ∥σc(ε⋆) − σB∥2.
Five benchmarks: AFHQ cat→dog (Type 2), CelebA-HQ man→woman (Type 2), breast frozen→permanent (Type 1, histopathology), day→night relighting (Type 1), and virtual staining unstained→H&E (Type 1). All images at 256×256, all splits are patient/scene-disjoint and held out.
Compared against CycleGAN, CUT, an SDEdit-style baseline, UNSB, and EGSDE. Every method is retrained on the identical held-out splits at 256×256, taken at its best validation-FID checkpoint under a comparable budget.
EGSDE is restricted to two benchmarks due to its slow sampler.
AFHQ cat→dog: PRISM achieves the lowest FID (76.9 versus 107.5 for the next best method) and the lowest KID (27.0 vs 65.0)
among evaluated methods, also ranking first in DINOv2- and CLIP-space Fréchet distances. Three-seed retraining gave FID 77.4 ± 0.7.
CelebA-HQ man→woman: PRISM achieves the best scores (FID 90.6, KID 53.9), using an early-stopped operating point that balances target-domain appearance with source identity preservation.
Breast frozen→permanent: PRISM achieves the lowest Inception FID (51.8) and KID (24.7)
outperforming CUT (56.1), UNSB (65.5), CycleGAN (74.6), SDEdit (139.5), and EGSDE (171.5). It also yields the nuclei-count ratio closest to the ideal value (0.93)
versus EGSDE's 3.32, indicating it best combines realism with a near-ideal nuclei budget.
Day→night: PRISM achieves the lowest FID (85.9) and KID (7.6), improving over CUT, the strongest baseline, at FID 97.7 and KID 14.5.
Virtual staining: CycleGAN and PRISM obtain comparable distributional performance: CycleGAN achieves slightly lower FID (49.2 versus 50.8), while PRISM achieves the lower KID (23.6 versus 24.9).
The authors regard them as competitive rather than claiming an advantage.
LNG: On AFHQ, LNG replaces the stochastic corruption with a learned deterministic initialization... (FID 75.0 versus 75.3). Removing the distribution-validity penalty Lndist degrades FID to 80.0.
Either criterion alone can be misleading. A method may improve realism by altering excessive source content, whereas high source similarity may result from insufficient translation.
On breast, PRISM combines the lowest Inception FID with the nuclei-count ratio closest to the ideal value of 1.0.
On AFHQ, high source similarity can indicate under-translation, so LPIPS and paired DINO similarity are interpreted together with FID rather than as independent measures of translation quality.
Because τ affects sampling only through the initialization and transport gate, the predicted field can be replaced at inference by a gate constructed from an external spatial cue, without retraining.
For natural images, CLIPSeg relevance maps from prompts like eyes
or background
increase τ in selected regions. For histopathology, a StarDist nuclei detector provides the spatial cue. Varying the override strength provides a continuous trade-off between source preservation and target-domain realism from a single trained checkpoint.
The gate-granularity ablation shows Performance improves as the gate becomes more fine-grained,
with FID decreasing from 74.1 (global) to 56.0 (spatial) to 49.1 (per-feature). The DDτ ablation shows DDτ improves nuclei preservation with little change in FID
; replacing content-anchored corruption with isotropic noise substantially degrades structural preservation
; Freezing a distilled gate performs worse than joint optimization
; and removing the constrained correction produces a different failure mode in which source preservation is retained but target realism deteriorates.
The lean-core ablation shows the gated-transport mechanism accounts for most of the performance and that the additional preservation terms provide secondary stabilization.
Analysis shows features assigned larger gate values generally undergo smaller latent displacements, whereas features with smaller gate values receive larger updates,
strongest for breast and weaker for AFHQ. Together with other experiments, these results support the interpretation of the gate as a feature-dependent preservation signal rather than merely an auxiliary prediction.
The paper lists four main limitations: "(1) the gate is validated through ablation, displacement, and inference-time override analyses rather than against ground-truth spatial change masks; (2) the gated dynamics are an empirical design and do not inherit a formal distributional guarantee from the frozen flow... (3) the pathology evaluation relies on automated structural proxies and does not replace blinded pathologist assessment or clinical validation; (4) the evaluation is limited by finite-sample metrics, single-run estimates on some benchmarks, separate training for each domain pair, and the 256×256 latent-resolution setting."
"PRISM achieves the lowest Inception FID and KID on four benchmarks and competitive performance on the fifth. On histopathology, it also yields the nuclei-count ratio closest to the ideal value, indicating improved structural preservation while maintaining realistic translation. Future work includes
a learned, differentiable controller that allocates integration steps on a per-feature basis, jointly optimizing translation quality and computational cost."
Improvements for AI systems
I can improve unpaired image-translation AI systems by incorporating the following mechanisms from PRISM.
- Per-feature preservation gate instead of a global noise knob
The system learns a field τ ∈ [0,1] over every latent feature and spatial location. High-τ features stay at source values; low-τ features are transported toward the target.
What it can do: translate images where different regions need opposite treatments—e.g., change a cat to a dog while preserving the background and pose, or stain a tissue slide while keeping nuclear outlines intact—without forcing one global strength on the whole image.
- Distribution-informed gate supervision (DDτ)
The gate is trained to match a robust prior computed from each feature’s standardized distance to the target-domain feature statistics: patches already close to the target are preserved, distant patches are changed.
What it can do: decide automatically where and how much to translate per image, without requiring paired data or ground-truth change masks.
- Inference-time gate override from external cues
Because the gate is only used at initialization and during integration, it can be replaced at inference with masks from CLIPSeg, object detectors, nuclei detectors, or text prompts, with no retraining.
What it can do: let users specify “keep the eyes” or “translate only the background” after training, and smoothly trade off preservation versus target realism from one checkpoint.
- Task-matched initialization
Use AdaIN-based content-anchored noise for structure-preserving tasks (histopathology, relighting) and partially anchored isotropic noise for structure-changing tasks (cat→dog, man→woman), with a blending coefficient α.
What it can do: address both translation regimes in a single framework—maintaining layout for structure-preserving tasks while allowing larger shape/appearance changes when needed.
- Stable non-adversarial training with constrained residual correction
Freeze a pretrained flow-matching model and add a small, zero-initialized, norm-constrained DiT that corrects off-manifold drift without overriding the base flow. Training uses MMD, feature losses, and gate losses—no GAN or PatchGAN.
What it can do: avoid adversarial training instability and mode collapse while still producing realistic target-domain outputs.
- Local k-NN target matching with DINOv2 features
Instead of one global target distribution, the system uses each source image’s k nearest real targets in DINOv2 space and minimizes a per-sample Mahalanobis distance to their statistics.
What it can do: translate each specific input toward the nearest plausible target sub-region, improving realism and avoiding generic or saturated translations for atypical inputs.
- Diagnosis-consistency regularization for medical translation
Stage 1 is conditioned on metadata (organ, diagnosis, grade), and Stage 2 includes a frozen diagnosis-classifier loss to keep the translated image’s diagnostic label consistent with the source.
What it can do: perform tasks like virtual staining or frozen→permanent section conversion while preserving pathologically meaningful structures and labels, not just appearance.
- Learned deterministic initialization (LNG)
A small generator produces one deterministic noise vector per input, with a distribution-validity penalty, making translation a reproducible one-to-one map rather than a stochastic draw.
What it can do: produce consistent, repeatable outputs for the same input image across runs—important for clinical or scientific reproducibility—while still matching the target distribution.
- Gate-driven per-feature compute allocation
The wake-up timing τ determines which features need early, full integration budget; preserved features can use fewer function evaluations.
What it can do: when combined with a future learned controller, allocate computation preferentially to regions that need translation, reducing inference cost for mostly-preserved images while retaining quality on high-change regions.
- Better realism–faithfulness evaluation
Use paired structural proxies (e.g., nuclei-count ratios, LPIPS with FID interpretation, DINO/CLIP Fréchet distances) rather than FID alone.
What it can do: allow the system and its users to detect under-translation versus hallucination, and choose an operating point that balances realism with source preservation instead of optimizing a single misleading metric.
Sources
- One-Step Image Translation with Text-to-Image Models
- Revisiting the state-space model of unawareness
- Classifier-Free Diffusion Guidance
- Hierarchical Sparse Attention Framework for Computationally Efficient Classification of Biological Cells
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models