Current fake audio detection algorithms have achieved promising performances on most datasets. However, their performance may be significantly degraded when dealing with audio of a different dataset. The orthogonal weight modification to overcome catastrophic forgetting does not consider the similarity of genuine audio across different datasets. To overcome this limitation, we propose a continual learning algorithm for fake audio detection to overcome catastrophic forgetting, called Regularized Adaptive Weight Modification (RAWM). When fine-tuning a detection network, our approach adaptively computes the direction of weight modification according to the ratio of genuine utterances and fake utterances. The adaptive modification direction ensures the network can effectively detect fake audio on the new dataset while preserving its knowledge of old model, thus mitigating catastrophic forgetting. In addition, genuine audio collected from quite different acoustic conditions may skew their feature distribution, so we introduce a regularization constraint to force the network to remember the old distribution in this regard. Our method can easily be generalized to related fields, like speech emotion recognition. We also evaluate our approach across multiple datasets and obtain a significant performance improvement on cross-dataset experiments.
翻译:当前假音频检测算法在大多数数据集上已取得可观性能。然而,当处理不同数据集的音频时,其性能可能显著下降。现有克服灾难性遗忘的正交权重修改方法未考虑不同数据集中真实音频的相似性。为克服这一局限,我们提出一种用于假音频检测的持续学习算法——正则化自适应权重修改(RAWM),以克服灾难性遗忘。在微调检测网络时,我们的方法根据真实语音与虚假语音的比例自适应计算权重修改方向。这种自适应修改方向确保网络能在保留旧模型知识的同时有效检测新数据集上的假音频,从而缓解灾难性遗忘。此外,来源自完全不同声学条件的真实音频可能导致其特征分布偏移,因此我们引入正则化约束强制网络记住该领域的旧分布。我们的方法可轻松推广至相关领域,如语音情感识别。我们还在多个数据集上评估了该方法,并在跨数据集实验中获得了显著的性能提升。