Generalisation -- the ability of a model to perform well on unseen data -- is crucial for building reliable deep fake detectors. However, recent studies have shown that the current audio deep fake models fall short of this desideratum. In this paper we show that pretrained self-supervised representations followed by a simple logistic regression classifier achieve strong generalisation capabilities, reducing the equal error rate from 30% to 8% on the newly introduced In-the-Wild dataset. Importantly, this approach also produces considerably better calibrated models when compared to previous approaches. This means that we can trust our model's predictions more and use these for downstream tasks, such as uncertainty estimation. In particular, we show that the entropy of the estimated probabilities provides a reliable way of rejecting uncertain samples and further improving the accuracy.
翻译:泛化能力——模型在未见数据上表现良好的能力——对于构建可靠的深度伪造检测器至关重要。然而,近期研究表明,当前音频深度伪造模型尚不能满足这一要求。本文证明,采用预训练的自监督表示结合简单的逻辑回归分类器,可具备强大的泛化能力,在最新引入的In-the-Wild数据集上,将等错误率从30%降至8%。尤为重要的是,相较于先前方法,该方法生成的模型校准性能显著提升。这意味着我们能够更信任模型的预测结果,并将其用于下游任务(如不确定性估计)。具体而言,我们证明估计概率的熵为拒绝不确定样本及进一步提升准确率提供了可靠途径。