Zero-shot sketch-based image retrieval (ZS-SBIR) is challenging due to the cross-domain nature of sketches and photos, as well as the semantic gap between seen and unseen image distributions. Previous methods fine-tune pre-trained models with various side information and learning strategies to learn a compact feature space that (\romannumeral1) is shared between the sketch and photo domains and (\romannumeral2) bridges seen and unseen classes. However, these efforts are inadequate in adapting domains and transferring knowledge from seen to unseen classes. In this paper, we present an effective \emph{``Adapt and Align''} approach to address the key challenges. Specifically, we insert simple and lightweight domain adapters to learn new abstract concepts of the sketch domain and improve cross-domain representation capabilities. Inspired by recent advances in image-text foundation models (\textit{e.g.}, CLIP) on zero-shot scenarios, we explicitly align the learned image embedding with a more semantic text embedding to achieve the desired knowledge transfer from seen to unseen classes. Extensive experiments on three benchmark datasets and two popular backbones demonstrate the superiority of our method in terms of retrieval accuracy and flexibility.
翻译:零样本手绘草图图像检索(ZS-SBIR)因草图和照片的跨域特性以及可见与不可见图分布之间的语义鸿沟而具有挑战性。以往方法通过不同辅助信息和学习策略微调预训练模型,以学习一个紧凑的特征空间,该空间:(Ⅰ)在草图和照片域之间共享;(Ⅱ)桥接可见和不可见类别。然而,这些方法在域适应以及将知识从可见类迁移至不可见类方面仍存在不足。本文提出一种有效的“调整与对齐”(Adapt and Align)方法以应对关键挑战。具体而言,我们插入简单轻量的域适配器来学习草图域的新抽象概念,并提升跨域表征能力。受近期图像-文本基础模型(如CLIP)在零样本场景中进展的启发,我们将学习到的图像嵌入与更具语义性的文本嵌入显式对齐,从而实现从可见类到不可见类的知识迁移。在三个基准数据集和两种主流骨干网络上的大量实验表明,本方法在检索精度和灵活性方面具有优越性。