Recent advances in vision-language learning have achieved notable success on complete-information question-answering datasets through the integration of extensive world knowledge. Yet, most models operate passively, responding to questions based on pre-stored knowledge. In stark contrast, humans possess the ability to actively explore, accumulate, and reason using both newfound and existing information to tackle incomplete-information questions. In response to this gap, we introduce $Conan$, an interactive open-world environment devised for the assessment of active reasoning. $Conan$ facilitates active exploration and promotes multi-round abductive inference, reminiscent of rich, open-world settings like Minecraft. Diverging from previous works that lean primarily on single-round deduction via instruction following, $Conan$ compels agents to actively interact with their surroundings, amalgamating new evidence with prior knowledge to elucidate events from incomplete observations. Our analysis on $Conan$ underscores the shortcomings of contemporary state-of-the-art models in active exploration and understanding complex scenarios. Additionally, we explore Abduction from Deduction, where agents harness Bayesian rules to recast the challenge of abduction as a deductive process. Through $Conan$, we aim to galvanize advancements in active reasoning and set the stage for the next generation of artificial intelligence agents adept at dynamically engaging in environments.
翻译:近期视觉-语言学习领域通过整合广泛的世界知识,在处理信息完整型问答数据集方面取得了显著成功。然而,大多数模型仍处于被动运行状态,仅依赖预存知识回答问题。与此形成鲜明对比的是,人类具备主动探索、积累并运用新信息与既有知识解决信息残缺问题的能力。为填补这一空白,我们提出$Conan$——一个用于评估主动推理能力的交互式开放世界环境。$Conan$框架鼓励主动探索并促进多轮溯因推理,其特性类似于《我的世界》这类丰富的开放世界场景。与先前主要依赖指令遵循进行单轮演绎推理的研究不同,$Conan$迫使智能体主动与环境交互,通过整合新证据与先验知识,从残缺信息中推导事件成因。我们在$Conan$上的分析揭示了当前最先进模型在主动探索与理解复杂场景方面的明显不足。此外,我们探究了"从演绎到溯因"机制——智能体利用贝叶斯规则将溯因挑战转化为演绎过程。通过$Conan$,我们期望推动主动推理领域的突破,为新时代能够动态适应环境的人工智能代理奠定基础。