We introduce Bongard-OpenWorld, a new benchmark for evaluating real-world few-shot reasoning for machine vision. It originates from the classical Bongard Problems (BPs): Given two sets of images (positive and negative), the model needs to identify the set that query images belong to by inducing the visual concepts, which is exclusively depicted by images from the positive set. Our benchmark inherits the few-shot concept induction of the original BPs while adding the two novel layers of challenge: 1) open-world free-form concepts, as the visual concepts in Bongard-OpenWorld are unique compositions of terms from an open vocabulary, ranging from object categories to abstract visual attributes and commonsense factual knowledge; 2) real-world images, as opposed to the synthetic diagrams used by many counterparts. In our exploration, Bongard-OpenWorld already imposes a significant challenge to current few-shot reasoning algorithms. We further investigate to which extent the recently introduced Large Language Models (LLMs) and Vision-Language Models (VLMs) can solve our task, by directly probing VLMs, and combining VLMs and LLMs in an interactive reasoning scheme. We even conceived a neuro-symbolic reasoning approach that reconciles LLMs & VLMs with logical reasoning to emulate the human problem-solving process for Bongard Problems. However, none of these approaches manage to close the human-machine gap, as the best learner achieves 64% accuracy while human participants easily reach 91%. We hope Bongard-OpenWorld can help us better understand the limitations of current visual intelligence and facilitate future research on visual agents with stronger few-shot visual reasoning capabilities.
翻译:我们提出了Bongard-OpenWorld,这是一个用于评估机器视觉在现实世界中进行少样本推理能力的新基准。它源于经典的Bongard问题(BPs):给定两组图像(正例和负例),模型需要通过归纳视觉概念(正例集特有)来判定查询图像所属的集合。我们的基准继承了原始BPs的少样本概念归纳特性,同时新增了两个挑战层面:1)开放世界的自由形态概念——Bongard-OpenWorld中的视觉概念是由开放词汇表中的术语组合而成的独特复合体,涵盖从物体类别到抽象视觉属性及常识事实性知识;2)真实世界图像——区别于许多同类基准使用的合成示意图。在探索中,Bongard-OpenWorld已对现有少样本推理算法构成重大挑战。我们进一步研究了近期推出的大型语言模型(LLMs)和视觉语言模型(VLMs)在多大程度上能解决我们的任务,具体采用了直接探测VLMs,以及在交互式推理方案中结合VLMs和LLMs的方法。我们甚至构想了一种神经符号推理方法,该方法将LLMs和VLMs与逻辑推理相结合,以模拟人类解决Bongard问题的过程。然而,这些方法均未能弥合人机差距:最优学习器的准确率仅为64%,而人类参与者轻松达到91%的准确率。我们希望Bongard-OpenWorld能帮助我们更好地理解当前视觉智能的局限性,并推动未来具备更强少样本视觉推理能力的视觉智能体研究。