Large multimodal models trained on natural documents, which interleave images and text, outperform models trained on image-text pairs on various multimodal benchmarks that require reasoning over one or multiple images to generate a text. However, the datasets used to train these models have not been released, and the collection process has not been fully specified. We introduce the OBELISC dataset, an open web-scale filtered dataset of interleaved image-text documents comprising 141 million web pages extracted from Common Crawl, 353 million associated images, and 115 billion text tokens. We describe the dataset creation process, present comprehensive filtering rules, and provide an analysis of the dataset's content. To show the viability of OBELISC, we train an 80 billion parameters vision and language model on the dataset and obtain competitive performance on various multimodal benchmarks. We release the code to reproduce the dataset along with the dataset itself.
翻译:在自然文档(即图像与文本交错排列的文档)上训练的大型多模态模型,在需要基于一张或多张图像生成文本的各种多模态基准测试中,其性能优于在图像-文本对数据上训练的模型。然而,用于训练这些模型的数据集尚未发布,且数据收集过程也未被完整说明。我们介绍OBELISC数据集,这是一个开放的网络规模过滤交错图像-文本文档数据集,包含从Common Crawl中提取的1.41亿个网页、3.53亿张关联图像以及1150亿个文本词元。我们描述了数据集的创建过程,提出了全面的过滤规则,并对数据集内容进行了分析。为验证OBELISC的可行性,我们基于该数据集训练了一个800亿参数的视觉语言模型,并在多项多模态基准测试中取得了具有竞争力的性能。我们公开了用于复现该数据集的代码以及数据集本身。