Purpose: To analyze the demographic and imaging characteristics associated with increased risk of failure for abnormality classification in screening mammograms. Materials and Methods: This retrospective study used data from the Emory BrEast Imaging Dataset (EMBED) which includes mammograms from 115,931 patients imaged at Emory University Healthcare between 2013 to 2020. Clinical and imaging data includes Breast Imaging Reporting and Data System (BI-RADS) assessment, region of interest coordinates for abnormalities, imaging features, pathologic outcomes, and patient demographics. Multiple deep learning models were developed to distinguish between patches of abnormal tissue and randomly selected patches of normal tissue from the screening mammograms. We assessed model performance overall and within subgroups defined by age, race, pathologic outcome, and imaging characteristics to evaluate reasons for misclassifications. Results: On a test set size of 5,810 studies (13,390 patches), a ResNet152V2 model trained to classify normal versus abnormal tissue patches achieved an accuracy of 92.6% (95% CI = 92.0-93.2%), and area under the receiver operative characteristics curve 0.975 (95% CI = 0.972-0.978). Imaging characteristics associated with higher misclassifications of images include higher tissue densities (risk ratio [RR]=1.649; p=.010, BI-RADS density C and RR=2.026; p=.003, BI-RADS density D), and presence of architectural distortion (RR=1.026; p<.001). Conclusion: Even though deep learning models for abnormality classification can perform well in screening mammography, we demonstrate certain imaging features that result in worse model performance. This is the first such work to systematically evaluate breast abnormality classification by various subgroups and better-informed developers and end-users of population subgroups which are likely to experience biased model performance.
翻译:目的:分析在筛查性乳腺X线摄影中与异常分类失败风险增加相关的人口统计学和影像学特征。材料与方法:本回顾性研究使用了埃默里乳腺影像数据集(EMBED)中的数据,该数据集包含2013年至2020年间在埃默里大学医疗系统成像的115,931名患者的乳腺X线摄影图像。临床和影像数据包括乳腺影像报告和数据系统(BI-RADS)评估、异常区域坐标、影像特征、病理结果和患者人口统计学信息。研究开发了多个深度学习模型,用于区分筛查性乳腺X线摄影图像中异常组织斑块与随机选取的正常组织斑块。我们评估了模型整体性能以及按年龄、种族、病理结果和影像特征划分的亚组性能,以分析误分类的原因。结果:在包含5,810次检查(13,390个斑块)的测试集上,用于区分正常与异常组织斑块的ResNet152V2模型达到了92.6%的准确率(95%置信区间=92.0-93.2%),受试者工作特征曲线下面积为0.975(95%置信区间=0.972-0.978)。与图像更高误分类率相关的影像特征包括更高的组织密度(风险比[RR]=1.649;p=0.010,BI-RADS密度C级;RR=2.026;p=0.003,BI-RADS密度D级)以及存在结构扭曲(RR=1.026;p<0.001)。结论:尽管用于异常分类的深度学习模型在筛查性乳腺X线摄影中表现良好,但我们证明了某些影像特征会导致模型性能下降。这是首次系统评估不同亚组中乳腺异常分类的研究,为开发者和终端用户提供了关于可能面临模型性能偏差的人口亚组的更充分信息。