Relation-focused cross-modal information retrieval focuses on retrieving information based on relations expressed in user queries, and it is particularly important in information retrieval applications and next-generation search engines. To date, CLIP (Contrastive Language-Image Pre-training) achieved state-of-the-art performance in cross-modal learning tasks due to its efficient learning of visual concepts from natural language supervision. However, CLIP learns visual representations from natural language at a global level without the capability of focusing on image-object relations. This paper proposes a novel CLIP-based network for Relation Reasoning, CLIP-RR, that tackles relation-focused cross-modal information retrieval. The proposed network utilises CLIP to leverage its pre-trained knowledge, and it additionally comprises two main parts: (1) extends the capabilities of CLIP to extract and reason with object relations in images; and (2) aggregates the reasoned results for predicting the similarity scores between images and descriptions. Experiments were carried out by applying the proposed network to relation-focused cross-modal information retrieval tasks on the RefCOCOg, CLEVR, and Flickr30K datasets. The results revealed that the proposed network outperformed various other state-of-the-art networks including CLIP, VSE$\infty$, and VSRN++ on both image-to-text and text-to-image cross-modal information retrieval tasks.
翻译:关系聚焦跨模态信息检索侧重于基于用户查询中表达的关系来检索信息,在信息检索应用和下一代搜索引擎中尤为重要。迄今为止,CLIP(对比语言-图像预训练)由于能从自然语言监督中高效学习视觉概念,在跨模态学习任务中取得了最先进的性能。然而,CLIP在全局层面上从自然语言学习视觉表征,缺乏聚焦于图像-对象关系的能力。本文提出了一种新颖的基于CLIP的关系推理网络CLIP-RR,用于解决关系聚焦跨模态信息检索问题。该网络利用CLIP的预训练知识,并额外包含两个主要部分:(1)扩展CLIP的能力,以提取和推理图像中的对象关系;(2)聚合推理结果,用于预测图像与描述之间的相似度分数。通过将该网络应用于RefCOCOg、CLEVR和Flickr30K数据集上的关系聚焦跨模态信息检索任务进行了实验。结果表明,该网络在图像到文本和文本到图像的跨模态信息检索任务上均优于包括CLIP、VSE∞和VSRN++在内的多种其他最先进网络。