Continual Test-Time Adaptation (CTTA) generalizes conventional Test-Time Adaptation (TTA) by assuming that the target domain is dynamic over time rather than stationary. In this paper, we explore Multi-Modal Continual Test-Time Adaptation (MM-CTTA) as a new extension of CTTA for 3D semantic segmentation. The key to MM-CTTA is to adaptively attend to the reliable modality while avoiding catastrophic forgetting during continual domain shifts, which is out of the capability of previous TTA or CTTA methods. To fulfill this gap, we propose an MM-CTTA method called Continual Cross-Modal Adaptive Clustering (CoMAC) that addresses this task from two perspectives. On one hand, we propose an adaptive dual-stage mechanism to generate reliable cross-modal predictions by attending to the reliable modality based on the class-wise feature-centroid distance in the latent space. On the other hand, to perform test-time adaptation without catastrophic forgetting, we design class-wise momentum queues that capture confident target features for adaptation while stochastically restoring pseudo-source features to revisit source knowledge. We further introduce two new benchmarks to facilitate the exploration of MM-CTTA in the future. Our experimental results show that our method achieves state-of-the-art performance on both benchmarks.
翻译:连续测试时自适应(CTTA)将传统测试时自适应(TTA)推广至目标域随时间动态变化而非静态的场景。本文探索了多模态连续测试时自适应(MM-CTTA)作为CTTA在三维语义分割中的新扩展。MM-CTTA的关键在于自适应地关注可靠模态,同时在连续域偏移过程中避免灾难性遗忘,这是以往TTA或CTTA方法无法做到的。为填补这一空白,我们提出一种名为连续跨模态自适应聚类(CoMAC)的MM-CTTA方法,从两个角度解决该任务。一方面,我们提出一种自适应双阶段机制,通过基于潜在空间中类别级特征质心距离来关注可靠模态,从而生成可靠的跨模态预测。另一方面,为在无灾难性遗忘的情况下进行测试时自适应,我们设计了类别级动量队列,该队列在自适应过程中捕获置信的目标特征,同时随机恢复伪源特征以重新激活源域知识。此外,我们引入两个新基准以促进未来对MM-CTTA的探索。实验结果表明,我们的方法在两个基准上均达到了最先进的性能。