The surge of interest towards Multi-modal Large Language Models (MLLMs), e.g., GPT-4V(ision) from OpenAI, has marked a significant trend in both academia and industry. They endow Large Language Models (LLMs) with powerful capabilities in visual understanding, enabling them to tackle diverse multi-modal tasks. Very recently, Google released Gemini, its newest and most capable MLLM built from the ground up for multi-modality. In light of the superior reasoning capabilities, can Gemini challenge GPT-4V's leading position in multi-modal learning? In this paper, we present a preliminary exploration of Gemini Pro's visual understanding proficiency, which comprehensively covers four domains: fundamental perception, advanced cognition, challenging vision tasks, and various expert capacities. We compare Gemini Pro with the state-of-the-art GPT-4V to evaluate its upper limits, along with the latest open-sourced MLLM, Sphinx, which reveals the gap between manual efforts and black-box systems. The qualitative samples indicate that, while GPT-4V and Gemini showcase different answering styles and preferences, they can exhibit comparable visual reasoning capabilities, and Sphinx still trails behind them concerning domain generalizability. Specifically, GPT-4V tends to elaborate detailed explanations and intermediate steps, and Gemini prefers to output a direct and concise answer. The quantitative evaluation on the popular MME benchmark also demonstrates the potential of Gemini to be a strong challenger to GPT-4V. Our early investigation of Gemini also observes some common issues of MLLMs, indicating that there still remains a considerable distance towards artificial general intelligence. Our project for tracking the progress of MLLM is released at https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models.
翻译:多模态大语言模型(MLLMs)的兴起,例如OpenAI开发的GPT-4V(ision),已成为学术界和工业界的重要趋势。这类模型赋予大语言模型(LLMs)强大的视觉理解能力,使其能够处理多样化的多模态任务。近期,谷歌发布了Gemini——其最新且功能最强大的多模态大语言模型,该模型从底层设计即为多模态而生。凭借卓越的推理能力,Gemini能否挑战GPT-4V在多模态学习领域的领先地位?本文对Gemini Pro的视觉理解能力进行了初步探索,全面覆盖四个领域:基础感知、高级认知、挑战性视觉任务以及多种专家能力。我们将Gemini Pro与最先进的GPT-4V进行对比以评估其上限,同时与最新开源的多模态大语言模型Sphinx进行比较,揭示人工构建系统与黑盒系统之间的差距。定性样本表明,GPT-4V与Gemini虽展现出不同的回答风格与偏好,但二者在视觉推理能力上可达到相近水平,而Sphinx在领域泛化性上仍逊色于它们。具体而言,GPT-4V倾向于提供详细的解释和中间步骤,而Gemini则偏好输出直接简洁的答案。基于流行多模态评估基准MME的定量评估也证明了Gemini具备成为GPT-4V强有力挑战者的潜力。我们对Gemini的早期探索还发现了MLLMs的一些常见问题,表明距离通用人工智能仍有相当距离。我们用于追踪MLLM进展的项目已发布在https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models。