Following the successful debut of polyp detection and characterization, more advanced automation tools are being developed for colonoscopy. The new automation tasks, such as quality metrics or report generation, require understanding of the procedure flow that includes activities, events, anatomical landmarks, etc. In this work we present a method for automatic semantic parsing of colonoscopy videos. The method uses a novel DL multi-label temporal segmentation model trained in supervised and unsupervised regimes. We evaluate the accuracy of the method on a test set of over 300 annotated colonoscopy videos, and use ablation to explore the relative importance of various method's components.
翻译:在息肉检测与特征识别成功应用后,结肠镜领域正开发更先进的自动化工具。新的自动化任务(如质量指标计算或报告生成)需理解包含活动、事件、解剖标志等要素的检查流程。本研究提出一种结肠镜视频自动语义解析方法,该方法采用新型深度学习多标签时序分割模型,通过监督与非监督联合训练策略进行优化。我们在包含300余段已标注结肠镜视频的测试集上评估了方法精度,并通过消融实验探究各组件对方法性能的相对贡献。