We introduce ViLPAct, a novel vision-language benchmark for human activity planning. It is designed for a task where embodied AI agents can reason and forecast future actions of humans based on video clips about their initial activities and intents in text. The dataset consists of 2.9k videos from \charades extended with intents via crowdsourcing, a multi-choice question test set, and four strong baselines. One of the baselines implements a neurosymbolic approach based on a multi-modal knowledge base (MKB), while the other ones are deep generative models adapted from recent state-of-the-art (SOTA) methods. According to our extensive experiments, the key challenges are compositional generalization and effective use of information from both modalities.
翻译:我们提出了ViLPAct,一个用于人类活动规划的新型视觉语言基准。它专为具身智能体能够基于视频片段中人类初始活动与文本意图进行推理并预测其未来动作的任务而设计。该数据集包含来自Charades的2.9k个视频片段(通过众包扩展了意图信息)、一个多项选择题测试集以及四个强基线模型。其中一个基线基于多模态知识库(MKB)实现了神经符号方法,其余基线则改编自近期最先进(SOTA)的深度生成模型。根据我们的广泛实验,关键挑战在于组合泛化能力及对两种模态信息的有效利用。