Screening is critical for prevention and early detection of cervical cancer but it is time-consuming and laborious. Supervised deep convolutional neural networks have been developed to automate pap smear screening and the results are promising. However, the interest in using only normal samples to train deep neural networks has increased owing to class imbalance problems and high-labeling costs that are both prevalent in healthcare. In this study, we introduce a method to learn explainable deep cervical cell representations for pap smear cytology images based on one class classification using variational autoencoders. Findings demonstrate that a score can be calculated for cell abnormality without training models with abnormal samples and localize abnormality to interpret our results with a novel metric based on absolute difference in cross entropy in agglomerative clustering. The best model that discriminates squamous cell carcinoma (SCC) from normals gives 0.908 +- 0.003 area under operating characteristic curve (AUC) and one that discriminates high-grade epithelial lesion (HSIL) 0.920 +- 0.002 AUC. Compared to other clustering methods, our method enhances the V-measure and yields higher homogeneity scores, which more effectively isolate different abnormality regions, aiding in the interpretation of our results. Evaluation using in-house and additional open dataset show that our model can discriminate abnormality without the need of additional training of deep models.
翻译:筛查对于宫颈癌的预防和早期发现至关重要,但其过程耗时费力。目前已开发出基于监督式深度卷积神经网络的方法用于自动化巴氏涂片筛查,并取得了良好的效果。然而,由于医疗领域中普遍存在的类别不平衡问题和高标注成本,仅使用正常样本训练深度神经网络的研究兴趣日益增加。在本研究中,我们提出了一种基于变分自编码器的一类分类方法,用于学习巴氏涂片细胞学图像的可解释深度宫颈细胞表征。研究结果表明,无需使用异常样本训练模型即可计算细胞异常评分,并且我们通过一种基于凝聚聚类中交叉熵绝对差值的新型度量方法定位异常区域,从而解释结果。在区分鳞状细胞癌(SCC)与正常样本的最佳模型中,受试者工作特征曲线下面积(AUC)达到0.908 ± 0.003;而区分高级别上皮内病变(HSIL)的模型AUC达到0.920 ± 0.002。与其他聚类方法相比,我们的方法提高了V-measure指标并获得了更高的同质性评分,从而更有效地分离不同异常区域,有助于结果的解释。基于内部数据集和额外公开数据集的评估表明,我们的模型无需深度模型额外训练即可区分异常。