The remarkable multimodal capabilities demonstrated by OpenAI's GPT-4 have sparked significant interest in the development of multimodal Large Language Models (LLMs). A primary research objective of such models is to align visual and textual modalities effectively while comprehending human instructions. Current methodologies often rely on annotations derived from benchmark datasets to construct image-dialogue datasets for training purposes, akin to instruction tuning in LLMs. However, these datasets often exhibit domain bias, potentially constraining the generative capabilities of the models. In an effort to mitigate these limitations, we propose a novel data collection methodology that synchronously synthesizes images and dialogues for visual instruction tuning. This approach harnesses the power of generative models, marrying the abilities of ChatGPT and text-to-image generative models to yield a diverse and controllable dataset with varied image content. This not only provides greater flexibility compared to existing methodologies but also significantly enhances several model capabilities. Our research includes comprehensive experiments conducted on various datasets using the open-source LLAVA model as a testbed for our proposed pipeline. Our results underscore marked enhancements across more than ten commonly assessed capabilities,
翻译:OpenAI的GPT-4所展现出的卓越多模态能力,引发了人们对多模态大语言模型发展的浓厚兴趣。此类模型的首要研究目标是在有效理解人类指令的同时,实现视觉与文本模态的对齐。当前方法通常依赖于从基准数据集中获取的标注,以构建用于训练的图像-对话数据集,这与大语言模型中的指令微调类似。然而,这些数据集往往存在领域偏差,可能限制模型的生成能力。为了缓解这些局限,我们提出了一种新颖的数据收集方法,该方法同步合成图像和对话以用于视觉指令微调。这一方法利用了生成模型的能力,将ChatGPT和文本到图像生成模型相结合,生成具有多样化图像内容的可控数据集。这不仅比现有方法提供了更大的灵活性,还显著提升了模型的多种能力。我们的研究包含了在多个数据集上进行的全面实验,使用开源LLAVA模型作为我们所提管线的测试平台。我们的结果强调了在超过十项常被评估的能力上均取得了显著提升,