Manual labeling of gestures in robot-assisted surgery is labor intensive, prone to errors, and requires expertise or training. We propose a method for automated and explainable generation of gesture transcripts that leverages the abundance of data for image segmentation. Surgical context is detected using segmentation masks by examining the distances and intersections between the tools and objects. Next, context labels are translated into gesture transcripts using knowledge-based Finite State Machine (FSM) and data-driven Long Short Term Memory (LSTM) models. We evaluate the performance of each stage of our method by comparing the results with the ground truth segmentation masks, the consensus context labels, and the gesture labels in the JIGSAWS dataset. Our results show that our segmentation models achieve state-of-the-art performance in recognizing needle and thread in Suturing and we can automatically detect important surgical states with high agreement with crowd-sourced labels (e.g., contact between graspers and objects in Suturing). We also find that the FSM models are more robust to poor segmentation and labeling performance than LSTMs. Our proposed method can significantly shorten the gesture labeling process (~2.8 times).
翻译:在机器人辅助手术中,手势标签的人工标注过程劳动强度大、易出错且需要专业知识或培训。我们提出一种自动化且可解释的手势转录生成方法,该方法利用了图像分割数据的丰富性。通过分析工具与物体之间的距离和交集,使用分割掩码检测手术情境。随后,基于知识的有穷状态机(FSM)和基于数据的长短期记忆(LSTM)模型,将情境标签转换对手势转录。通过将各阶段结果与JIGSAWS数据集中的地面真值分割掩码、共识情境标签和手势标签进行对比,我们评估了方法的性能。结果表明,我们的分割模型在缝合任务中识别针与线方面达到最优性能,且能自动检测重要手术状态,与众包标签高度一致(例如缝合中夹钳与物体的接触)。此外,我们发现FSM模型对分割和标签性能的鲁棒性优于LSTM。所提方法可大幅缩短手势标注流程(约2.8倍)。