Despite the successes of recent developments in visual AI, different shortcomings still exist; from missing exact logical reasoning, to abstract generalization abilities, to understanding complex and noisy scenes. Unfortunately, existing benchmarks, were not designed to capture more than a few of these aspects. Whereas deep learning datasets focus on visually complex data but simple visual reasoning tasks, inductive logic datasets involve complex logical learning tasks, however, lack the visual component. To address this, we propose the visual logical learning dataset, V-LoL, that seamlessly combines visual and logical challenges. Notably, we introduce the first instantiation of V-LoL, V-LoL-Trains, -- a visual rendition of a classic benchmark in symbolic AI, the Michalski train problem. By incorporating intricate visual scenes and flexible logical reasoning tasks within a versatile framework, V-LoL-Trains provides a platform for investigating a wide range of visual logical learning challenges. We evaluate a variety of AI systems including traditional symbolic AI, neural AI, as well as neuro-symbolic AI. Our evaluations demonstrate that even state-of-the-art AI faces difficulties in dealing with visual logical learning challenges, highlighting unique advantages and limitations specific to each methodology. Overall, V-LoL opens up new avenues for understanding and enhancing current abilities in visual logical learning for AI systems.
翻译:尽管近期视觉人工智能的发展取得了诸多成功,但仍存在不同短板:从精准逻辑推理的缺失,到抽象泛化能力的不足,再到理解复杂且含噪场景的局限。遗憾的是,现有基准测试在设计上仅能捕捉其中少数几个方面。深度学习数据集侧重于视觉上复杂的数据但推理任务简单,而归纳逻辑数据集虽涉及复杂的逻辑学习任务,却缺乏视觉成分。为解决这一问题,我们提出了视觉逻辑学习数据集V-LoL,该数据集无缝融合了视觉与逻辑挑战。值得注意的是,我们引入了V-LoL的首个实例——V-LoL-Trains,这是符号人工智能经典基准测试(Michalski列车问题)的视觉化呈现。通过在一个通用框架中集成精细的视觉场景与灵活的逻辑推理任务,V-LoL-Trains为探索广泛的视觉逻辑学习挑战提供了平台。我们评估了多种人工智能系统,包括传统符号人工智能、神经网络人工智能以及神经符号人工智能。评估结果表明,即使是当前最先进的人工智能在处理视觉逻辑学习挑战时仍面临困难,并凸显了每种方法特有的优势与局限。总体而言,V-LoL为理解并增强当前人工智能系统在视觉逻辑学习方面的能力开辟了新途径。