Affordance knowledge is a fundamental aspect of commonsense knowledge. Recent findings indicate that world knowledge emerges through large-scale self-supervised pretraining, motivating our exploration of acquiring affordance knowledge from the visual domain. To this end, we augment an existing instructional video resource to create the new Causal Action-Effect (CAE) dataset and design two novel pretraining tasks -- Masked Action Modeling (MAM) and Masked Effect Modeling (MEM) -- promoting the acquisition of two affordance properties in models: behavior and entity equivalence, respectively. We empirically demonstrate the effectiveness of our proposed methods in learning affordance properties. Furthermore, we show that a model pretrained on both tasks outperforms a strong image-based visual-linguistic foundation model (FLAVA) as well as pure linguistic models on a zero-shot physical reasoning probing task.
翻译:可供性知识是常识知识的基本组成部分。最新研究表明,通过大规模自监督预训练可涌现世界知识,这促使我们探索从视觉领域获取可供性知识。为此,我们对现有教学视频资源进行扩充,创建了新的因果动作-效应(CAE)数据集,并设计了两种新型预训练任务——掩码动作建模(MAM)和掩码效应建模(MEM)——分别促进模型获取可供性的两种属性:行为等价性与实体等价性。我们通过实验证明了所提方法在学习可供性属性方面的有效性。进一步研究表明,在零样本物理推理探测任务上,同时预训练这两种任务的模型超越了强大的基于图像的视觉-语言基础模型(FLAVA)以及纯语言模型。