Early diagnosis of Type 2 Diabetes Mellitus (T2DM) is crucial to enable timely therapeutic interventions and lifestyle modifications. As the time available for clinical office visits shortens and medical imaging data become more widely available, patient image data could be used to opportunistically identify patients for additional T2DM diagnostic workup by physicians. We investigated whether image-derived phenotypic data could be leveraged in tabular learning classifier models to predict T2DM risk in an automated fashion to flag high-risk patients without the need for additional blood laboratory measurements. In contrast to traditional binary classifiers, we leverage neural networks and decision tree models to represent patient data as 'SynthA1c' latent variables, which mimic blood hemoglobin A1c empirical lab measurements, that achieve sensitivities as high as 87.6%. To evaluate how SynthA1c models may generalize to other patient populations, we introduce a novel generalizable metric that uses vanilla data augmentation techniques to predict model performance on input out-of-domain covariates. We show that image-derived phenotypes and physical examination data together can accurately predict diabetes risk as a means of opportunistic risk stratification enabled by artificial intelligence and medical imaging. Our code is available at https://github.com/allisonjchae/DMT2RiskAssessment.
翻译:2型糖尿病的早期诊断对于及时实施治疗干预和生活方式调整至关重要。随着临床门诊时间缩短及医学影像数据日益普及,患者影像数据可被用于机会性识别需进一步接受T2DM诊断检查的患者。本研究探究了是否可利用影像衍生表型数据,通过表格学习分类器模型以自动化方式预测T2DM风险,从而无需额外血液实验室检测即可标记高风险患者。与传统二元分类器不同,我们利用神经网络和决策树模型将患者数据表征为“SynthA1c”潜变量,该变量模拟血液糖化血红蛋白经验实验室测量值,实现了高达87.6%的灵敏度。为评估SynthA1c模型对其他患者群体的泛化能力,我们提出了一种新型可泛化度量指标,该指标使用原始数据增强技术预测模型在输入域外协变量上的性能。研究表明,影像衍生表型与体格检查数据相结合,可准确预测糖尿病风险,从而实现基于人工智能与医学影像的机会性风险分层。我们的代码已开源:https://github.com/allisonjchae/DMT2RiskAssessment。