In this report, we introduce NICE (New frontiers for zero-shot Image Captioning Evaluation) project and share the results and outcomes of 2023 challenge. This project is designed to challenge the computer vision community to develop robust image captioning models that advance the state-of-the-art both in terms of accuracy and fairness. Through the challenge, the image captioning models were tested using a new evaluation dataset that includes a large variety of visual concepts from many domains. There was no specific training data provided for the challenge, and therefore the challenge entries were required to adapt to new types of image descriptions that had not been seen during training. This report includes information on the newly proposed NICE dataset, evaluation methods, challenge results, and technical details of top-ranking entries. We expect that the outcomes of the challenge will contribute to the improvement of AI models on various vision-language tasks.
翻译:本报告介绍了NICE(零样本图像描述评估新前沿)项目,并分享了2023年挑战赛的结果与成果。该项目旨在激励计算机视觉领域开发鲁棒的图像描述模型,以在准确性和公平性两方面推动技术前沿。挑战赛采用包含跨领域大量视觉概念的新评估数据集测试图像描述模型。由于未提供特定训练数据,参赛模型需适应训练阶段未见的新型图像描述。本报告涵盖新提出的NICE数据集、评估方法、挑战赛结果及排名前列提交方案的技术细节。我们期待挑战赛的成果有助于提升AI模型在各类视觉-语言任务中的表现。