Current captioning datasets, focus on object-centric captions, describing the visible objects in the image, often ending up stating the obvious (for humans), e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Language models to recognize the visual content, they lack in expressing trivial abstract concepts, e.g. "people having a picnic". Such concepts are licensed by human's personal experience and contribute to forming common sense assumptions. We present the High-Level Dataset; a dataset extending 14997 images of the COCO dataset with 134973 human-annotated (high-level) abstract captions collected along three axes: scenes, actions and rationales. We describe and release such dataset and we show how it can be used to assess models' multimodal grounding of abstract concepts and enrich models' visio-lingusitic representations. Moreover, we describe potential tasks enabled by this dataset involving high- and low-level concepts interactions.
翻译:当前的图像描述数据集主要关注以物体为核心的描述,即描述图像中可见的物体,往往陈述对人类而言显而易见的内容,例如“人们在公园里吃饭”。虽然这些数据集有助于评估视觉与语言模型识别视觉内容的能力,但它们在表达类似“人们正在野餐”这样的简单抽象概念方面存在不足。这些概念源于人类的个人经验,并有助于形成常识性假设。我们提出了高层数据集(High-Level Dataset);该数据集扩展了COCO数据集的14997张图像,新增了由人工标注的、沿着场景、动作和理由三个轴收集的134973条(高层)抽象描述。我们对该数据集进行了描述并予以公开,同时展示了如何利用它来评估模型对抽象概念的多模态锚定能力,并丰富模型的视觉-语言表征。此外,我们还描述了由该数据集支持的、涉及高层与低层概念交互的潜在任务。