Phacoemulsification cataract surgery (PCS) is a routine procedure conducted using a surgical microscope, heavily reliant on the skill of the ophthalmologist. While existing PCS guidance systems extract valuable information from surgical microscopic videos to enhance intraoperative proficiency, they suffer from non-phasespecific guidance, leading to redundant visual information. In this study, our major contribution is the development of a novel phase-specific augmented reality (AR) guidance system, which offers tailored AR information corresponding to the recognized surgical phase. Leveraging the inherent quasi-standardized nature of PCS procedures, we propose a two-stage surgical microscopic video recognition network. In the first stage, we implement a multi-task learning structure to segment the surgical limbus region and extract limbus region-focused spatial feature for each frame. In the second stage, we propose the long-short spatiotemporal aggregation transformer (LS-SAT) network to model local fine-grained and global temporal relationships, and combine the extracted spatial features to recognize the current surgical phase. Additionally, we collaborate closely with ophthalmologists to design AR visual cues by utilizing techniques such as limbus ellipse fitting and regional restricted normal cross-correlation rotation computation. We evaluated the network on publicly available and in-house datasets, with comparison results demonstrating its superior performance compared to related works. Ablation results further validated the effectiveness of the limbus region-focused spatial feature extractor and the combination of temporal features. Furthermore, the developed system was evaluated in a clinical setup, with results indicating remarkable accuracy and real-time performance. underscoring its potential for clinical applications.
翻译:超声乳化白内障手术(PCS)是在手术显微镜下开展的常规术式,高度依赖眼科医师的操作技能。现有PCS导航系统虽能从显微手术视频中提取有价值信息以提升术中效率,但其非相位感知的导航方式会导致视觉信息冗余。本研究的主要贡献在于开发了一种新型相位感知增强现实(AR)导航系统,该系统可根据识别的手术相位提供定制化AR信息。利用PCS流程固有的准标准化特性,我们提出两阶段显微手术视频识别网络:第一阶段采用多任务学习结构分割手术角膜缘区域,并提取每帧图像的角膜缘区域聚焦空间特征;第二阶段提出长-短时空聚合Transformer(LS-SAT)网络,通过建模局部细粒度与全局时序关系,结合所提取空间特征识别当前手术相位。此外,我们与眼科医师密切协作,采用角膜缘椭圆拟合与区域受限归一化互相关旋转计算等技术设计AR视觉提示。在公开数据集与内部数据集上的评估表明,该网络性能显著优于现有相关研究。消融实验进一步验证了角膜缘区域聚焦空间特征提取器及时序特征组合的有效性。临床环境下的系统测试结果显示其兼具卓越精度与实时性能,凸显了该系统的临床转化潜力。