Building machines that can reason about physical events and their causal relationships is crucial for flexible interaction with the physical world. However, most existing physical and causal reasoning benchmarks are exclusively based on synthetically generated events and synthetic natural language descriptions of causal relationships. This design brings up two issues. First, there is a lack of diversity in both event types and natural language descriptions; second, causal relationships based on manually-defined heuristics are different from human judgments. To address both shortcomings, we present the CLEVRER-Humans benchmark, a video reasoning dataset for causal judgment of physical events with human labels. We employ two techniques to improve data collection efficiency: first, a novel iterative event cloze task to elicit a new representation of events in videos, which we term Causal Event Graphs (CEGs); second, a data augmentation technique based on neural language generative models. We convert the collected CEGs into questions and answers to be consistent with prior work. Finally, we study a collection of baseline approaches for CLEVRER-Humans question-answering, highlighting the great challenges set forth by our benchmark.
翻译:构建能够推理物理事件及其因果关系的机器,对于实现与物理世界的灵活交互至关重要。然而,现有的大多数物理与因果推理基准完全基于合成生成的事件及其因果关系的合成自然语言描述。这种设计带来了两个问题:首先,事件类型和自然语言描述的多样性不足;其次,基于人工定义启发式规则的因果关系与人类的判断存在差异。为弥补这两方面的不足,我们提出了CLEVRER-Humans基准——一个包含人类标注的物理事件因果判断视频推理数据集。我们采用两种技术提高数据收集效率:其一,一种新颖的迭代事件完形填空任务,用于提取视频中事件的新表示,我们称之为因果事件图(CEGs);其二,一种基于神经语言生成模型的数据增强技术。我们将收集到的CEGs转换为与先前工作一致的问答形式。最后,我们研究了一系列针对CLEVRER-Humans问答任务的基线方法,突显了本基准带来的重大挑战。