Mode Collapse Is Cheap to Detect: A Ground-Truth-Free Pre-Flight Check for Neural Samplers

arXiv:2609.26272 · cs.LG · Submitted 2026-08-20 · Read on arXiv

cs.LG

Submitted: 2026-08-20

Updated: 2026-08-20

License: http://creativecommons.org/licenses/by/4.0/

The gist: Neural samplers are trained against an unnormalised target π=e-E with no samples from π, which leaves the practitioner with no way to tell whether an expensive training run has silently dropped

Terminology

Abstract

Neural samplers are trained against an unnormalised target π=e-E with no samples from π, which leaves the practitioner with no way to tell whether an expensive training run has silently dropped part of the target. The diagnostics in common use are computed from the model's own draws and are therefore confined to the model's support: we exhibit a sampler whose self-normalised effective sample size is 0.99 while it misses 87% of the target mass. We argue that detecting missing mass is a strictly easier problem than sampling it: detection needs one point per missed basin plus a local curvature estimate, whereas correction needs the sampler retrained. We turn this into a pre-flight check that consumes a few percent of the sampler's own training budget and uses only E, grad E and grad squared E. On Gaussian-mixture, Many-Well and rotated anisotropic Many-Well targets with exactly computable ground truth, the check estimates the missing mass to within 10-3 at 2.7% of training cost, where a tuned annealed SMC reference needs 70 -- 280% of training cost to do worse. It also applies unchanged to a controlled-SDE sampler that has no tractable density, where ESS and the ELBO cannot be formed at all. The estimator carries a self-diagnostic that, without ground truth, is conservative in the safe direction: across 60 configurations it clears 16, of which 15 are accurate to 10-2 or better. We are explicit about what this does and does not license: the check cheaply produces evidence of missing mass, and sometimes evidence that the search has stabilised, but it cannot certify a run, and its thresholds are heuristic. We then map the boundary of the method on a real physical landscape, LJ-13, and report where it fails and why.

Sources

Related papers