Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection
eess.AS, cs.CL, cs.SD
Submitted: 2026-05-30
Updated: 2026-09-14
Comments: 16 pages, 5 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: We address the problem of out-of-distribution (OOD) detection for target observations embedded in a subspace of the high dimensional data space.
Terminology
Abstract
We address the problem of out-of-distribution (OOD) detection for target observations embedded in a subspace of the high dimensional data space. Using continuous normalizing flows (CNFs), we propose a Lagrangian sub-flow (LSF) framework designed to isolate and estimate the density for the relevant components in the representation and using the remaining components as context. Through experimentation with models for speech synthesis, we show that CNFs, similarly to other deep generative models (DGMs), are susceptible to the "likelihood paradox", where high likelihood is erroneously assigned to OOD samples. This is attributed to the inductive bias of DGMs that prioritize low-level structural details over high-level semantic coherence. To mitigate this phenomenon, we propose a number of geometric diagnostic signals based on the velocity field over the sub-flow trajectory. Based on these signals, we design metrics for the challenging task of zero-shot phoneme-level mispronunciation detection. Finally, we demonstrate the superiority of these metrics compared to likelihood-based methods on a real-world mispronunciation detection benchmark.
Sources
- Equivariant flow-based sampling for lattice gauge theory
- F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions