Instruction tuning has significantly advanced large language models (LLMs) such as ChatGPT, enabling them to align with human instructions across diverse tasks. However, progress in open vision-language models (VLMs) has been limited due to the scarcity of high-quality instruction datasets. To tackle this challenge and promote research in the vision-language field, we introduce the Multi-Modal, Multilingual Instruction Tuning (M$^3$IT) dataset, designed to optimize VLM alignment with human instructions. Our M$^3$IT dataset comprises 40 carefully curated datasets, including 2.4 million instances and 400 manually written task instructions, reformatted into a vision-to-text structure. Key tasks are translated into 80 languages with an advanced translation system, ensuring broader accessibility. M$^3$IT surpasses previous datasets regarding task coverage, instruction number and instance scale. Moreover, we develop Ying-VLM, a VLM model trained on our M$^3$IT dataset, showcasing its potential to answer complex questions requiring world knowledge, generalize to unseen video tasks, and comprehend unseen instructions in Chinese. To encourage further research, we have open-sourced both the dataset and trained models.
翻译:指令微调显著提升了ChatGPT等大型语言模型(LLMs)的能力,使其能够在不同任务中与人类指令对齐。然而,由于高质量指令数据集的匮乏,开放视觉语言模型(VLMs)的进展仍受限。为应对这一挑战并推动视觉语言领域研究,我们提出了多模态多语言指令微调(M$^3$IT)数据集,旨在优化VLM与人类指令的对齐。我们的M$^3$IT数据集包含40个精心整理的数据集,共计240万条实例和400条人工编写的任务指令,并统一重构为视觉到文本(vision-to-text)的结构。关键任务通过先进的翻译系统被翻译为80种语言,确保更广泛的适用性。在任务覆盖范围、指令数量及实例规模上,M$^3$IT均超越了以往数据集。此外,我们基于M$^3$IT数据集训练了Ying-VLM模型,展示了其在回答需要世界知识的复杂问题、泛化至未见视频任务以及理解中文未见过指令方面的潜力。为促进进一步研究,我们已开源该数据集及训练模型。