Vision-language datasets are vital for both text-to-image (T2I) and image-to-text (I2T) research. However, current datasets lack descriptions with fine-grained detail that would allow for richer associations to be learned by models. To fill the gap, we introduce Descriptions of Connected and Contrasting Images (DOCCI), a dataset with long, human-annotated English descriptions for 15k images that were taken, curated and donated by a single researcher intent on capturing key challenges such as spatial relations, counting, text rendering, world knowledge, and more. We instruct human annotators to create comprehensive descriptions for each image; these average 136 words in length and are crafted to clearly distinguish each image from those that are related or similar. Each description is highly compositional and typically encompasses multiple challenges. Through both quantitative and qualitative analyses, we demonstrate that DOCCI serves as an effective training resource for image-to-text generation -- a PaLI 5B model finetuned on DOCCI shows equal or superior results compared to highly-performant larger models like LLaVA-1.5 7B and InstructBLIP 7B. Furthermore, we show that DOCCI is a useful testbed for text-to-image generation, highlighting the limitations of current text-to-image models in capturing long descriptions and fine details.
翻译:视觉语言数据集对于文本到图像(T2I)和图像到文本(I2T)研究至关重要。然而,当前数据集缺乏包含细粒度细节的描述,这使得模型难以学习更丰富的关联。为填补这一空白,我们引入了关联与对比图像描述数据集(DOCCI),该数据集包含由单一研究者拍摄、整理并捐赠的15,000张图像的长篇人工标注英文描述,旨在应对空间关系、计数、文本渲染、世界知识等关键挑战。我们指导标注员为每张图像创建详尽的描述,这些描述平均长度为136个单词,并经过精心设计以明确区分每张图像与其相关或相似图像。每条描述高度组合化,通常涵盖多重挑战。通过定量与定性分析,我们证明DOCCI可作为图像到文本生成的有效训练资源——基于DOCCI微调的PaLI 5B模型展现出与高性能大型模型(如LLaVA-1.5 7B和InstructBLIP 7B)相当甚至更优的结果。此外,我们表明DOCCI是文本到图像生成的有用测试平台,揭示了当前文本到图像模型在捕捉长描述与细粒度细节方面的局限性。