Teaching Visual Question Answering (VQA) models to abstain from unanswerable questions is indispensable for building a trustworthy AI system. Existing studies, though have explored various aspects of VQA, yet marginally ignored this particular attribute. This paper aims to bridge the research gap by contributing a comprehensive dataset, called UNK-VQA. The dataset is specifically designed to address the challenge of questions that can be unanswerable. To this end, we first augment the existing data via deliberate perturbations on either the image or question. In specific, we carefully ensure that the question-image semantics remain close to the original unperturbed distribution. By means of this, the identification of unanswerable questions becomes challenging, setting our dataset apart from others that involve mere image replacement. We then extensively evaluate the zero- and few-shot performance of several emerging multi-modal large models and discover significant limitations of them when applied to our dataset. Additionally, we also propose a straightforward method to tackle these unanswerable questions. This dataset, we believe, will serve as a valuable benchmark for enhancing the abstention capability of VQA models, thereby leading to increased trustworthiness of AI systems.
翻译:教会视觉问答(VQA)模型在面对无法回答的问题时主动放弃回答,是构建可信人工智能系统的必要条件。现有研究虽然探索了VQA的多个方面,却鲜少关注这一特性。本文旨在通过构建一个名为UNK-VQA的综合性数据集来填补这一研究空白。该数据集专门针对可能无法回答的问题挑战而设计。为此,我们首先通过对图像或问题进行有意的扰动来增强现有数据。具体而言,我们精心确保扰动后的问题与图像语义仍接近原始未扰动分布。如此一来,识别无法回答问题的难度显著提升,使得我们的数据集区别于仅涉及图像替换的其他数据集。随后,我们广泛评估了多种新兴多模态大模型的零样本和少样本性能,发现在我们的数据集上这些模型存在显著局限性。此外,我们还提出了一种直接方法来处理这些无法回答的问题。我们相信,该数据集将成为提升VQA模型拒绝回答能力的宝贵基准,从而增强AI系统的可信度。