The increased interest in diffusion models has opened up opportunities for advancements in generative text modeling. These models can produce impressive images when given a well-crafted prompt, but creating a powerful or meaningful prompt can be hit-or-miss. To address this, we have created a large-scale dataset that is derived and synthesized from real prompts and indexed with popular image-text datasets such as MS-COCO and Flickr. We have also implemented stages that gradually reduce context and increase complexity, which will further enhance the output due to the complex annotations created. The dataset, called MTTN, includes over 2.4 million sentences divided into 5 stages, resulting in a total of over 12 million pairs, and a vocabulary of over 300,000 unique words, providing ample variation. The original 2.4 million pairs are designed to reflect the way language is used on the internet globally, making the dataset more robust for any model trained on it.
翻译:扩散模型日益增长的兴趣为生成式文本建模的进步带来了机遇。这些模型在给定精心设计的提示时能生成令人印象深刻的图像,但创建有效或有意义的提示却可能成败难料。为解决这一问题,我们构建了一个大规模数据集,该数据集从真实提示中推导并合成而来,并与MS-COCO、Flickr等主流图像-文本数据集建立索引。我们还实现了逐步减少上下文并增加复杂度的多阶段处理流程,借助所创建的复杂标注进一步提升输出质量。该数据集名为MTTN,包含超过240万条句子,划分为5个阶段,总计生成超过1200万个配对,词汇量达30余万独特单词,提供了丰富的变体。原始240万配对的设计旨在反映全球互联网语言的使用方式,从而使任何基于该数据集训练的模型更具鲁棒性。