Accurate and robust tracking and reconstruction of the surgical scene is a critical enabling technology toward autonomous robotic surgery. Existing algorithms for 3D perception in surgery mainly rely on geometric information, while we propose to also leverage semantic information inferred from the endoscopic video using image segmentation algorithms. In this paper, we present a novel, comprehensive surgical perception framework, Semantic-SuPer, that integrates geometric and semantic information to facilitate data association, 3D reconstruction, and tracking of endoscopic scenes, benefiting downstream tasks like surgical navigation. The proposed framework is demonstrated on challenging endoscopic data with deforming tissue, showing its advantages over our baseline and several other state-of the-art approaches. Our code and dataset are available at https://github.com/ucsdarclab/Python-SuPer.
翻译:精准鲁棒的手术场景追踪与重建是实现自主机器人手术的关键使能技术。现有手术三维感知算法主要依赖几何信息,而本文提出利用图像分割算法从内窥镜视频中推断的语义信息。本文提出一种新型综合性手术感知框架——Semantic-SuPer,通过融合几何与语义信息,促进内窥镜场景的数据关联、三维重建与追踪,为手术导航等下游任务提供支撑。该框架在具有组织形变挑战性的内窥镜数据上进行了验证,结果表明其相较于基线方法及其他多种前沿方法具有显著优势。相关代码与数据集已开源至https://github.com/ucsdarclab/Python-SuPer。