This tutorial note summarizes the presentation on ``Large Multimodal Models: Towards Building and Surpassing Multimodal GPT-4'', a part of CVPR 2023 tutorial on ``Recent Advances in Vision Foundation Models''. The tutorial consists of three parts. We first introduce the background on recent GPT-like large models for vision-and-language modeling to motivate the research in instruction-tuned large multimodal models (LMMs). As a pre-requisite, we describe the basics of instruction-tuning in large language models, which is further extended to the multimodal space. Lastly, we illustrate how to build the minimum prototype of multimodal GPT-4 like models with the open-source resource, and review the recently emerged topics.
翻译:本教程笔记总结了题为“大型多模态模型:构建与超越多模态GPT-4”的报告内容,该报告是CVPR 2023“视觉基础模型最新进展”教程的一部分。教程包含三部分:首先,我们介绍近期类GPT大型模型在视觉与语言建模领域的背景,以启发对指令微调大型多模态模型(LMMs)的研究;作为前置知识,我们阐述大型语言模型中指令微调的基本原理,并将其进一步拓展至多模态空间;最后,我们演示如何利用开源资源构建类多模态GPT-4模型的最小原型,并回顾近期涌现的相关研究课题。