The hubness problem, in which hub embeddings are close to many unrelated examples, occurs often in high-dimensional embedding spaces and may pose a practical threat for purposes such as information retrieval and automatic evaluation metrics. In particular, since cross-modal similarity between text and images cannot be calculated by direct comparisons, such as string matching, cross-modal encoders that project different modalities into a shared space are helpful for various cross-modal applications, and thus, the existence of hubs may pose practical threats. To reveal the vulnerabilities of cross-modal encoders, we propose a method for identifying the hub embedding and its corresponding hub text. Experiments on image captioning evaluation in MSCOCO and nocaps along with image-to-text retrieval tasks in MSCOCO and Flickr30k showed that our method can identify a single hub text that unreasonably achieves comparable or higher similarity scores than human-written reference captions in many images, thereby revealing the vulnerabilities in cross-modal encoders.
翻译:中心性问题(hubness problem)是指中心嵌入(hub embeddings)与许多无关示例距离过近的现象,在高维嵌入空间中经常出现,并可能对信息检索和自动评估指标等应用构成实际威胁。特别是,由于文本与图像之间的跨模态相似性无法通过直接比较(如字符串匹配)来计算,因此将不同模态投影到共享空间的跨模态编码器在各种跨模态应用中十分有用,而中心性的存在可能带来实际威胁。为揭示跨模态编码器的脆弱性,我们提出了一种识别中心嵌入及其对应中心文本的方法。在MSCOCO和nocaps的图像描述评估任务,以及MSCOCO和Flickr30k的图像到文本检索任务上的实验表明,我们的方法能够识别出单个中心文本,该文本不合理地在多个图像中达到了与人工编写参考描述相当或更高的相似度得分,从而揭示了跨模态编码器的脆弱性。