Recent advances achieved by deep learning models rely on the independent and identically distributed assumption, hindering their applications in real-world scenarios with domain shifts. To tackle this issue, cross-domain learning aims at extracting domain-invariant knowledge to reduce the domain shift between training and testing data. However, in visual cross-domain learning, traditional methods concentrate solely on the image modality, disregarding the potential benefits of incorporating the text modality. In this work, we propose VLLaVO, combining Vision language models and Large Language models as Visual cross-dOmain learners. VLLaVO uses vision-language models to convert images into detailed textual descriptions. A large language model is then finetuned on textual descriptions of the source/target domain generated by a designed instruction template. Extensive experimental results under domain generalization and unsupervised domain adaptation settings demonstrate the effectiveness of the proposed method.
翻译:近年来深度学习模型取得的进展依赖于独立同分布假设,这限制了它们在存在域偏移的真实场景中的应用。为解决该问题,跨域学习旨在提取域不变知识以减小训练数据与测试数据之间的域偏移。然而,在视觉跨域学习中,传统方法仅关注图像模态,忽略了引入文本模态的潜在优势。本文提出VLLaVO,将视觉语言模型和大语言模型结合作为视觉跨域学习器。VLLaVO利用视觉语言模型将图像转换为详细的文本描述,然后通过设计的指令模板对源域/目标域的文本描述对大语言模型进行微调。在域泛化和无监督域适应设置下的大量实验结果表明了该方法的有效性。