Lifelong audio feature extraction involves learning new sound classes incrementally, which is essential for adapting to new data distributions over time. However, optimizing the model only on new data can lead to catastrophic forgetting of previously learned tasks, which undermines the model's ability to perform well over the long term. This paper introduces a new approach to continual audio representation learning called DeCoR. Unlike other methods that store previous data, features, or models, DeCoR indirectly distills knowledge from an earlier model to the latest by predicting quantization indices from a delayed codebook. We demonstrate that DeCoR improves acoustic scene classification accuracy and integrates well with continual self-supervised representation learning. Our approach introduces minimal storage and computation overhead, making it a lightweight and efficient solution for continual learning.
翻译:终身音频特征提取涉及增量式学习新声音类别,这对于随时间适应新的数据分布至关重要。然而,仅在新数据上优化模型可能导致先前学习任务的灾难性遗忘,从而损害模型长期的良好表现。本文提出一种名为DeCoR的持续音频表示学习新方法。与其他存储先前数据、特征或模型的方法不同,DeCoR通过从延迟码本中预测量化索引,间接将知识从早期模型提炼到最新模型。我们证明,DeCoR可提升声学场景分类的准确性,并能与持续自监督表示学习良好集成。我们的方法引入了极小的存储和计算开销,使其成为持续学习中轻量且高效的解决方案。