Mode Collapse Is Cheap to Detect: A Ground-Truth-Free Pre-Flight Check for Neural Samplers
cs.LG
Submitted: 2026-08-20
Updated: 2026-08-20
License: http://creativecommons.org/licenses/by/4.0/
The gist: Neural samplers are trained against an unnormalised target π=e-E with no samples from π, which leaves the practitioner with no way to tell whether an expensive training run has silently dropped
Terminology
Abstract
Neural samplers are trained against an unnormalised target π=e-E with no samples from π, which leaves the practitioner with no way to tell whether an expensive training run has silently dropped part of the target. The diagnostics in common use are computed from the model's own draws and are therefore confined to the model's support: we exhibit a sampler whose self-normalised effective sample size is 0.99 while it misses 87% of the target mass. We argue that detecting missing mass is a strictly easier problem than sampling it: detection needs one point per missed basin plus a local curvature estimate, whereas correction needs the sampler retrained. We turn this into a pre-flight check that consumes a few percent of the sampler's own training budget and uses only E, grad E and grad squared E. On Gaussian-mixture, Many-Well and rotated anisotropic Many-Well targets with exactly computable ground truth, the check estimates the missing mass to within 10-3 at 2.7% of training cost, where a tuned annealed SMC reference needs 70 -- 280% of training cost to do worse. It also applies unchanged to a controlled-SDE sampler that has no tractable density, where ESS and the ELBO cannot be formed at all. The estimator carries a self-diagnostic that, without ground truth, is conservative in the safe direction: across 60 configurations it clears 16, of which 15 are accurate to 10-2 or better. We are explicit about what this does and does not license: the check cheaply produces evidence of missing mass, and sometimes evidence that the search has stabilised, but it cannot certify a run, and its thresholds are heuristic. We then map the boundary of the method on a real physical landscape, LJ-13, and report where it fails and why.
Sources
- Fisher meets Feynman: score-based variational inference with a product of experts
- Conditional Diffusion Sampling
- Diffusion models recover accurate mixture weights despite score function insensitivity
- Flow Sampling: Learning to Sample from Unnormalized Densities via Denoising Conditional Processes
- Large-scale Score-based Variational Posterior Inference for Bayesian Deep Neural Networks
- FlowVAT: Normalizing Flow Variational Inference with Affine-Invariant Tempering
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks