Language has been useful in extending the vision encoder to data from diverse distributions without empirical discovery in training domains. However, as the image description is mostly at coarse-grained level and ignores visual details, the resulted embeddings are still ineffective in overcoming complexity of domains at inference time. We present a self-supervision framework WIDIn, Wording Images for Domain-Invariant representation, to disentangle discriminative visual representation, by only leveraging data in a single domain and without any test prior. Specifically, for each image, we first estimate the language embedding with fine-grained alignment, which can be consequently used to adaptively identify and then remove domain-specific counterpart from the raw visual embedding. WIDIn can be applied to both pretrained vision-language models like CLIP, and separately trained uni-modal models like MoCo and BERT. Experimental studies on three domain generalization datasets demonstrate the effectiveness of our approach.
翻译:语言在将视觉编码器扩展至不同分布的数据方面具有重要作用,而无需在训练域中进行经验性探索。然而,由于图像描述通常停留在粗粒度层面且忽略了视觉细节,由此产生的嵌入表示在推理时仍难以有效应对域的复杂性。本文提出一种自监督框架WIDIn(通过图像描述实现域不变表示),仅利用单域数据且无需任何测试先验,以解耦出具有判别性的视觉表示。具体而言,对于每幅图像,我们首先通过细粒度对齐估计其语言嵌入,进而利用该嵌入自适应地识别并从原始视觉嵌入中移除域特异性成分。WIDIn可同时应用于CLIP等预训练视觉语言模型,以及MoCo和BERT等分别训练的单模态模型。在三个域泛化数据集上的实验研究验证了本方法的有效性。