3D patient body modeling is critical to the success of automated patient positioning for smart medical scanning and operating rooms. Existing CNN-based end-to-end patient modeling solutions typically require a) customized network designs demanding large amount of relevant training data, covering extensive realistic clinical scenarios (e.g., patient covered by sheets), which leads to suboptimal generalizability in practical deployment, b) expensive 3D human model annotations, i.e., requiring huge amount of manual effort, resulting in systems that scale poorly. To address these issues, we propose a generic modularized 3D patient modeling method consists of (a) a multi-modal keypoint detection module with attentive fusion for 2D patient joint localization, to learn complementary cross-modality patient body information, leading to improved keypoint localization robustness and generalizability in a wide variety of imaging (e.g., CT, MRI etc.) and clinical scenarios (e.g., heavy occlusions); and (b) a self-supervised 3D mesh regression module which does not require expensive 3D mesh parameter annotations to train, bringing immediate cost benefits for clinical deployment. We demonstrate the efficacy of the proposed method by extensive patient positioning experiments on both public and clinical data. Our evaluation results achieve superior patient positioning performance across various imaging modalities in real clinical scenarios.
翻译:三维患者身体建模对于实现智能医疗扫描和手术室中的自动化患者定位至关重要。现有的基于CNN的端到端患者建模方案通常需要:a) 定制的网络设计,要求大量相关训练数据覆盖广泛的现实临床场景(例如患者被床单覆盖),这导致实际部署中泛化能力欠佳;b) 昂贵的三维人体模型标注,即需要大量人工工作,导致系统可扩展性差。为解决这些问题,我们提出了一种通用的模块化三维患者建模方法,包括:(a) 基于注意力融合的多模态关键点检测模块,用于二维患者关节定位,学习跨模态互补的患者身体信息,从而在广泛的成像模态(如CT、MRI等)和临床场景(如严重遮挡)中提升关键点定位的鲁棒性和泛化能力;(b) 自监督三维网格回归模块,无需昂贵的三维网格参数标注即可训练,为临床部署带来直接成本效益。通过在公共和临床数据上进行广泛的患者定位实验,我们证明了所提出方法的有效性。评估结果在真实临床场景中展现了跨多种成像模态的优越患者定位性能。