Deep generative models trained by maximum likelihood remain very popular methods for reasoning about data probabilistically. However, it has been observed that they can assign higher likelihoods to out-of-distribution (OOD) data than in-distribution data, thus calling into question the meaning of these likelihood values. In this work we provide a novel perspective on this phenomenon, decomposing the average likelihood into a KL divergence term and an entropy term. We argue that the latter can explain the curious OOD behaviour mentioned above, suppressing likelihood values on datasets with higher entropy. Although our idea is simple, we have not seen it explored yet in the literature. This analysis provides further explanation for the success of OOD detection methods based on likelihood ratios, as the problematic entropy term cancels out in expectation. Finally, we discuss how this observation relates to recent success in OOD detection with manifold-supported models, for which the above decomposition does not hold directly.
翻译:通过最大似然训练的深度生成模型仍是概率数据推理的主流方法。然而,已有研究发现这类模型可能对分布外数据赋予比分布内数据更高的似然值,从而引发对似然值含义的质疑。本文提供了解析这一现象的新视角,将平均似然分解为KL散度项与熵项。我们认为后者可解释上述反常的OOD行为——即高熵数据集会抑制似然值。尽管该思想简洁明了,但文献中尚未对此展开探讨。该分析进一步解释了基于似然比的OOD检测方法为何有效,因为问题性熵项在期望中相互抵消。最后,我们讨论这一发现与近期基于流形支持模型的OOD检测成功案例间的关联——此类模型中上述分解无法直接成立。