Continual zero-shot learning involves learning seen classes incrementally while improving the ability to recognize unseen or yet-to-be-seen classes. It has a broad range of potential applications in real-world vision tasks, such as accelerating species discovery. However, in these scenarios, the changes in environmental conditions cause shifts in the presentation of captured images, which we refer to as domain shift, and adds complexity to the tasks. In this paper, we introduce Domain Aware Continual Zero-Shot Learning (DACZSL), a task that involves visually recognizing images of unseen categories in unseen domains continually. To address the challenges of DACZSL, we propose a Domain-Invariant Network (DIN). We empoly a dual network structure to learn factorized features to alleviate forgetting, where consists of a global shared net for domian-invirant and task-invariant features, and per-task private nets for task-specific features. Furthermore, we introduce a class-wise learnable prompt to obtain better class-level text representation, which enables zero-shot prediction of future unseen classes. To evaluate DACZSL, we introduce two benchmarks: DomainNet-CZSL and iWildCam-CZSL. Our results show that DIN significantly outperforms existing baselines and achieves a new state-of-the-art.
翻译:持续零样本学习涉及增量式地学习已见类别,同时提升对未见或未来类别的识别能力。该技术在现实视觉任务中具有广泛的应用前景,例如加速物种发现。然而,在这些场景中,环境条件的变化会导致所捕获图像呈现方式的漂移(称为领域偏移),增加了任务的复杂性。本文提出领域感知的持续零样本学习(Domain Aware Continual Zero-Shot Learning, DACZSL),该任务要求持续识别来自未见领域的未见类别图像。为解决DACZSL的挑战,我们提出领域不变网络(Domain-Invariant Network, DIN)。采用双网络结构学习分解特征以缓解遗忘:其中包含一个全局共享网络用于提取领域不变与任务不变特征,以及各任务私有网络用于提取任务特定特征。此外,我们引入类别级可学习提示以获得更优的类别级文本表示,从而实现对未来未见类别的零样本预测。为评估DACZSL,我们构建了两个基准数据集:DomainNet-CZSL和iWildCam-CZSL。实验结果表明,DIN显著优于现有基线方法,达到了新的最优性能。