This work explores the problem of generating task graphs of real-world activities. Different from prior formulations, we consider a setting where text transcripts of instructional videos performing a real-world activity (e.g., making coffee) are provided and the goal is to identify the key steps relevant to the task as well as the dependency relationship between these key steps. We propose a novel task graph generation approach that combines the reasoning capabilities of instruction-tuned language models along with clustering and ranking components to generate accurate task graphs in a completely unsupervised manner. We show that the proposed approach generates more accurate task graphs compared to a supervised learning approach on tasks from the ProceL and CrossTask datasets.
翻译:本研究探索了从现实活动生成任务图的问题。与以往方法不同,我们考虑以下场景:给定执行现实活动(如制作咖啡)的教学视频文本转录,目标是识别与任务相关的关键步骤以及这些关键步骤之间的依赖关系。我们提出了一种新颖的任务图生成方法,该方法结合了指令调优语言模型的推理能力以及聚类和排序组件,以完全无监督的方式生成准确的任务图。实验表明,在ProceL和CrossTask数据集的任务上,所提方法相比监督学习方法能生成更准确的任务图。