This work focuses on anticipating long-term human actions, particularly using short video segments, which can speed up editing workflows through improved suggestions while fostering creativity by suggesting narratives. To this end, we imbue a transformer network with a symbolic knowledge graph for action anticipation in video segments by boosting certain aspects of the transformer's attention mechanism at run-time. Demonstrated on two benchmark datasets, Breakfast and 50Salads, our approach outperforms current state-of-the-art methods for long-term action anticipation using short video context by up to 9%.
翻译:本工作聚焦于预测人类长期动作,特别是利用短视频片段,通过改进建议来加速编辑流程,同时借助叙事提示激发创意。为此,我们将符号知识图谱融入Transformer网络,在运行时有针对性地增强其注意力机制的特定环节,以实现视频片段中的动作预测。在Breakfast与50Salads两个基准数据集上的实验表明,本方法在利用短视频上下文进行长期动作预测任务中,相较于当前最优方法提升了最多9%的性能。