Visual and linguistic concepts naturally organize themselves in a hierarchy, where a textual concept ``dog'' entails all images that contain dogs. Despite being intuitive, current large-scale vision and language models such as CLIP do not explicitly capture such hierarchy. We propose MERU, a contrastive model that yields hyperbolic representations of images and text. Hyperbolic spaces have suitable geometric properties to embed tree-like data, so MERU can better capture the underlying hierarchy in image-text data. Our results show that MERU learns a highly interpretable representation space while being competitive with CLIP's performance on multi-modal tasks like image classification and image-text retrieval.
翻译:视觉和语言概念天然地以层次结构组织,其中"狗"这一文本概念涵盖所有包含狗的图像。尽管直观易懂,但当前大规模视觉-语言模型(如CLIP)并未显式捕捉这种层次关系。我们提出MERU——一种生成图像与文本双曲表示的对比学习模型。双曲空间具有适合嵌入树状数据的几何特性,因此MERU能更有效地捕捉图像-文本数据中的潜在层次结构。实验结果表明,MERU在图像分类、图文检索等多模态任务中与CLIP性能相当的同时,学习到高度可解释的表示空间。