This paper does not present a novel method. Instead, it delves into an essential, yet must-know baseline in light of the latest advancements in Generative Artificial Intelligence (GenAI): the utilization of GPT-4 for visual understanding. Our study centers on the evaluation of GPT-4's linguistic and visual capabilities in zero-shot visual recognition tasks. Specifically, we explore the potential of its generated rich textual descriptions across various categories to enhance recognition performance without any training. Additionally, we evaluate its visual proficiency in directly recognizing diverse visual content. To achieve this, we conduct an extensive series of experiments, systematically quantifying the performance of GPT-4 across three modalities: images, videos, and point clouds. This comprehensive evaluation encompasses a total of 16 widely recognized benchmark datasets, providing top-1 and top-5 accuracy metrics. Our study reveals that leveraging GPT-4's advanced linguistic knowledge to generate rich descriptions markedly improves zero-shot recognition. In terms of visual proficiency, GPT-4V's average performance across 16 datasets sits roughly between the capabilities of OpenAI-CLIP's ViT-L and EVA-CLIP's ViT-E. We hope that this research will contribute valuable data points and experience for future studies. We release our code at https://github.com/whwu95/GPT4Vis.
翻译:本文并未提出新方法,而是基于生成式人工智能(GenAI)的最新进展,探讨一个关键且必须掌握的基础性问题:利用GPT-4进行视觉理解。研究聚焦于评估GPT-4在零样本视觉识别任务中的语言与视觉能力。具体而言,我们探索其生成的多类别丰富文本描述如何在不经训练的情况下提升识别性能,同时评估其直接识别多样视觉内容的视觉能力。为实现这一目标,我们开展了大量系统性实验,从图像、视频和点云三种模态定量衡量GPT-4的表现。评估涵盖共计16个广泛认可的基准数据集,并提供top-1与top-5准确率指标。研究表明,利用GPT-4的高级语言知识生成丰富描述可显著改进零样本识别效果;在视觉能力方面,GPT-4V在16个数据集上的平均性能介于OpenAI-CLIP的ViT-L与EVA-CLIP的ViT-E之间。我们期望本研究能为未来工作提供有价值的数据与经验。相关代码已发布于https://github.com/whwu95/GPT4Vis。