Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. While recent CIL advances have spurred significant interest across various modalities, the audio-visual setting remains underexplored. Furthermore, although foundational multimodal models like SAM-Audio encapsulate rich static priors, our empirical analysis reveals that these representations struggle in incremental settings. This work bridges this gap by integrating SAM-Audio's audio-visual priors into the CIL setting. Specifically, we leverage its dense audio and visual representations and employ a novel guided attention strategy where the audio features contextually guide the visual representations. To further mitigate catastrophic forgetting, we introduce dual-level distillation objectives at both the feature and logit levels. Extensive evaluations on audio-visual CIL benchmarks demonstrate that our approach consistently outperforms state-of-the-art methods.
翻译:类增量学习(CIL)旨在持续学习新类别而不遗忘先前获取的知识。尽管近期CIL的进展在多种模态中引发了广泛关注,但音视频场景仍未得到充分探索。此外,尽管如SAM-Audio等基础多模态模型封装了丰富的静态先验知识,我们的实证分析揭示这些表征在增量学习场景中存在困难。本研究通过将SAM-Audio的音视频先验知识整合至CIL框架来填补这一空白。具体而言,我们利用其密集的音频与视觉表征,并采用新颖的引导注意力机制——通过音频特征上下文引导视觉表征。为进一步缓解灾难性遗忘,我们引入特征级与logit级双层级蒸馏目标。在音视频CIL基准上的广泛评估表明,我们的方法始终优于现有最优方法。