In this work, we discuss building performant Multimodal Large Language Models (MLLMs). In particular, we study the importance of various architecture components and data choices. Through careful and comprehensive ablations of the image encoder, the vision language connector, and various pre-training data choices, we identified several crucial design lessons. For example, we demonstrate that for large-scale multimodal pre-training using a careful mix of image-caption, interleaved image-text, and text-only data is crucial for achieving state-of-the-art (SOTA) few-shot results across multiple benchmarks, compared to other published pre-training results. Further, we show that the image encoder together with image resolution and the image token count has substantial impact, while the vision-language connector design is of comparatively negligible importance. By scaling up the presented recipe, we build MM1, a family of multimodal models up to 30B parameters, consisting of both dense models and mixture-of-experts (MoE) variants, that are SOTA in pre-training metrics and achieve competitive performance after supervised fine-tuning on a range of established multimodal benchmarks. Thanks to large-scale pre-training, MM1 enjoys appealing properties such as enhanced in-context learning, and multi-image reasoning, enabling few-shot chain-of-thought prompting.
翻译:在本文中,我们探讨了构建高性能多模态大语言模型(MLLMs)的方法。特别地,我们研究了各类架构组件与数据选择的重要性。通过对图像编码器、视觉语言连接器以及多种预训练数据选择的细致全面消融实验,我们识别出若干关键设计经验。例如,我们证明,与其他已发表的预训练结果相比,大规模多模态预训练中采用图像-标题数据、交错图像-文本数据与纯文本数据的精心混合,对于在多个基准上实现最先进的少样本结果至关重要。此外,我们发现图像编码器及其分辨率与图像令牌数量具有显著影响,而视觉语言连接器的设计相比之下几乎可以忽略不计。通过扩展所提出的方案,我们构建了MM1,一个参数多达300亿的多模态模型家族,包含密集模型与混合专家(MoE)变体,其在预训练指标上达到最佳水平,并在经过监督微调后,在一系列成熟的多模态基准上展现出具有竞争力的性能。得益于大规模预训练,MM1展现出增强的上下文学习与多图像推理等吸引人的特性,使其支持少样本思维链提示。