Fine-Tuning Generative Models for Extreme Events via CVaR-Penalized Wasserstein Gradient Flows
University of Massachusetts Amherst · University of North Carolina at Chapel Hill
stat.ML, cs.LG
Submitted: 2026-08-12
Updated: 2026-10-01
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The paper proposes CVaR-GPA (CVaR-penalized Generative Particle Algorithm), a robust, tail-agnostic algorithm for fine-tuning pre-trained generative models to learn heavy-tailed distributions and
Terminology
Summary
The paper proposes CVaR-GPA (CVaR-penalized Generative Particle Algorithm), a robust, tail-agnostic algorithm for fine-tuning pre-trained generative models to learn heavy-tailed distributions and capture extreme events, requiring no prior knowledge or estimation of the target's tail characteristics. The method is defined as the Wasserstein gradient flow of a loss functional that penalizes the Lipschitz-regularized Kullback-Leibler (KL) divergence with a weighted, squared Conditional Value-at-Risk (CVaR) discrepancy term. The loss functional is given by:
F CVaR (Q; P tar) = D L KL (Q∥P tar) + λ (CVaRPα tar,g − CVaRQ,g α) squared
where D L KL denotes the Lipschitz-regularized KL divergence, CVaRQ,g α is the CVaR at the αth quantile of a non-negative risk function g under distribution Q, and λ > 0 is a hyperparameter controlling the relative weight of the squared CVaR difference and the Lipschitz-regularized KL divergence.
The paper states: "The Lipschitz-regularized KL divergence enables robust learning under minimal assumptions on the target distribution, while the CVaR penalty restores the velocity that otherwise vanishes prematurely in the under-sampled tails. The penalized flow admits a bounded but non-Lipschitz velocity field, which
departs from the Lipschitz transport maps of standard generators, which preserve the tail behavior of a light-tailed source, and enables transport toward heavier-tailed targets."
To define this flow on empirical measures, the authors derive the first-variation subgradients of CVaR from its Rockafellar-Uryasev representation, valid precisely where the classical density-based formula fails. The paper notes: "the variational derivative of CVaRQ,g α exists if and only if VaRQ,g α = VaRα Q,g... This condition fails for a given α whenever ΨQ,g (c) = α on a nonempty interval of c; for empirical measures, whose CDF ΨQ,g is a step function, such α always exist." Therefore, they use Clarke's generalized gradients to define variational subgradients.
The velocity field of the CVaR-penalized Wasserstein gradient flow is derived in Theorem 3.8. For each y ∈ T(Q), choosing the potential function ΦyQ induces the velocity field:
v y Q (x) = −∇x ϕ∗(x) + (2λ/(1−α)) ∆C(Q; P tar) ∇x g(x), for g(x) > y,
v y Q (x) = −∇x ϕ∗(x), for g(x) ≤ y,
where ϕ∗ is the maximizer of the variational representation of the Lipschitz-regularized KL divergence, and ∆C(Q; P tar) is the CVaR discrepancy. The paper emphasizes: "The CVaR penalization introduces an additional velocity component to the Lipschitz-regularized Wasserstein gradient flow in the tail region... Its velocity field is bounded, enabling stable learning, and not Lipschitz continuous."
The particle algorithm CVaR-GPA fine-tunes the output samples of any pre-trained model, without access to its architecture, and runs on an adaptive time horizon set by a kinetic-energy stopping criterion rather than a preset depth. The algorithm terminates when the empirical estimate of the kinetic energy Kk < ϵ for a given threshold ϵ, resulting in an algorithm whose architecture depth is implicitly determined by the target distribution P tar.
The paper evaluates CVaR-GPA on synthetic isotropic and anisotropic Student-t target distributions, Neal's funnel distribution, and the real-world high-dimensional Fama-French 25 portfolio dataset. The results show that CVaR-GPA dramatically improves global and tail accuracy on heavy-tailed targets over the pre-trained baseline.
Specifically:
-
For 2-dimensional isotropic Student-t distributions with tail indices ν ∈ 1, 1.2, 1.5, 1.8,
CVaR-GPA consistently reduces both the global L1 error EL1 and the tail error Etail of the pre-trained model across all tail indices,
demonstrating tail-agnostic capability. -
For the Fama-French 25 monthly portfolios dataset (25-dimensional anisotropic),
CVaR-GPA consistently decreases both the global L1 error and the tail error across all dimensions compared to the pre-trained model.
-
For Neal's funnel distribution,
CVaR-GPA improves both marginals by reducing the global L1 error and the tail error relative to the pre-trained model.
-
For a 5-dimensional anisotropic Student-t distribution with tail indices ν = (1, 1.5, 3, 10, 30), "CVaR-GPA consistently reduces both the global L1 error and the tail error for the heavy-tailed marginals with ν = (1, 1.5, 3, 10). For the nearly Gaussian marginal with ν = 30, however, both errors increase relative to the pre-trained model," indicating a limitation when there is extreme disparity between heavy-tailed and nearly Gaussian marginals.
The paper concludes: "We propose a tail-agnostic algorithm, CVaR-GPA, to fine-tune pre-trained models to learn multivariate heavy-tailed distributions with both anisotropic and isotropic tails. The CVaR penalization mitigates the effects of data scarcity in the tail region, remedying the premature saturation issue exhibited by pre-trained models." Future directions include using per-coordinate risk functions to adapt the CVaR-penalization to each marginal distribution, utilizing other spectral risk measures, and smoothing the function (g − y)+ at g(·) = y so that its first variational derivative is well-defined.
Improvements for AI systems
Improvements to AI Systems:
- Tail-Aware Generative Fine-Tuning Module
Integrate CVaR-GPA as a post-processing layer for any pre-trained generative model (e.g., diffusion models, GANs, VAEs) to automatically correct under-sampled tail regions without retraining the base model. The improved system can generate samples that faithfully capture extreme events (e.g., financial crashes, rare weather events) even when the training data has sparse tail observations.
- Adaptive Stopping Criterion for Inference Depth
Replace fixed-depth generation loops with a kinetic-energy-based stopping rule (as in CVaR-GPA). The improved system dynamically determines how many refinement steps to run per sample, reducing computational cost for simple inputs while allocating more steps for complex, heavy-tailed regions—enabling real-time adaptive inference.
- Robust Distribution Shift Handling
Use the Lipschitz-regularized KL + CVaR penalty as a fine-tuning objective for models deployed in non-stationary environments. The improved system can adapt to new target distributions with unknown tail behavior (e.g., shifting market regimes) without requiring explicit tail-index estimation, maintaining performance under distributional drift.
- Subgradient-Based Optimization for Non-Smooth Losses
Adopt Clarke’s generalized gradients (as derived for CVaR on empirical measures) to train models with loss functions that are non-differentiable at quantile boundaries. The improved system can optimize objectives involving CVaR, VaR, or other risk measures directly on finite samples, avoiding the failure modes of density-based formulas.
- Per-Coordinate Risk Calibration
Extend the CVaR penalty to per-coordinate risk functions (as suggested in future work). The improved system can independently tune tail behavior for each feature dimension, handling mixed distributions (e.g., some heavy-tailed, some Gaussian) without degrading performance on well-behaved marginals—solving the limitation observed with ν=30.
- Bounded Non-Lipschitz Velocity Fields for Stable Transport
Implement the derived velocity field (bounded but non-Lipschitz) in particle-based samplers or normalizing flows. The improved system can transport light-tailed source distributions to heavier-tailed targets without the instability of unbounded gradients, enabling stable training on extreme-value datasets.
- Tail-Agnostic Anomaly Detection
Repurpose the CVaR discrepancy term as a scoring function for detecting out-of-distribution samples. The improved system can flag rare events by measuring the gap between a sample’s CVaR under the model vs. the target, without assuming a specific tail shape—useful for fraud detection, network intrusion, or equipment failure prediction.
- Automatic Hyperparameter Selection via Tail Discrepancy
Use the squared CVaR difference as a self-tuning signal to adjust λ (the penalty weight) during training. The improved system can automatically increase penalization when tail errors dominate and decrease it when global accuracy saturates, removing manual tuning for each new dataset.
Sources
- Expected Shortfall as a Tool for Financial Risk Management
- Portfolio Optimization with Spectral Measures of Risk
- Flexible Tails for Normalizing Flows
- Spectral Normalization for Generative Adversarial Networks
- Heavy-Tailed Diffusion Models
- Efficient Tail-Aware Generative Optimization via Flow Model Fine-Tuning
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey