We address the problem of out-of-distribution (OOD) detection for target observations embedded in a subspace of the high dimensional data space. Using continuous normalizing flows (CNFs), we propose a Lagrangian sub-flow (LSF) framework designed to isolate and estimate the density for the relevant components in the representation and using the remaining components as context. Through experimentation with models for speech synthesis, we show that CNFs, similarly to other deep generative models (DGMs), are susceptible to the "likelihood paradox", where high likelihood is erroneously assigned to OOD samples. This is attributed to the inductive bias of DGMs that prioritize low-level structural details over high-level semantic coherence. To mitigate this phenomenon, we propose a number of geometric diagnostic signals based on the velocity field over the sub-flow trajectory. Based on these signals, we design metrics for the challenging task of zero-shot phoneme-level mispronunciation detection. Finally, we demonstrate the superiority of these metrics compared to likelihood-based methods on a real-world mispronunciation detection benchmark.
翻译:针对嵌入高维数据空间子空间中的目标观测,我们研究了分布外检测问题。基于连续归一化流(CNF),提出拉格朗日子流(LSF)框架,旨在分离并估计表示中相关分量的密度,同时将剩余分量作为上下文。通过对语音合成模型的实验表明,CNF与其他深度生成模型(DGM)类似,易受"似然悖论"影响——即对分布外样本错误赋予高似然值。该现象归因于DGM的归纳偏置:优先捕捉低层级结构细节而非高层级语义一致性。为缓解此问题,我们基于子流轨迹的速度场提出了若干几何诊断信号,并据此为零样本音素级错音检测这一挑战性任务设计度量指标。最终,在真实错音检测基准上验证了这些指标相较于基于似然方法具有更优性能。