Purpose - To characterise and assess the quality of published research evaluating artificial intelligence (AI) methods for ovarian cancer diagnosis or prognosis using histopathology data. Methods - A search of 5 sources was conducted up to 01/12/2022. The inclusion criteria required that research evaluated AI on histopathology images for diagnostic or prognostic inferences in ovarian cancer, including tubo-ovarian and peritoneal tumours. Reviews and non-English language articles were excluded. The risk of bias was assessed for every included model using PROBAST. Results - A total of 1434 research articles were identified, of which 36 were eligible for inclusion. These studies reported 62 models of interest, including 35 classifiers, 14 survival prediction models, 7 segmentation models, and 6 regression models. Models were developed using 1-1375 slides from 1-664 ovarian cancer patients. A wide array of outcomes were predicted, including overall survival (9/62), histological subtypes (7/62), stain quantity (6/62) and malignancy (5/62). Older studies used traditional machine learning (ML) models with hand-crafted features, while newer studies typically employed deep learning (DL) to automatically learn features and predict the outcome(s) of interest. All models were found to be at high or unclear risk of bias overall. Research was frequently limited by insufficient reporting, small sample sizes, and insufficient validation. Conclusion - Limited research has been conducted and none of the associated models have been demonstrated to be ready for real-world implementation. Recommendations are provided addressing underlying biases and flaws in study design, which should help inform higher-quality reproducible future research. Key aspects include more transparent and comprehensive reporting, and improved performance evaluation using cross-validation and external validations.
翻译:目的 - 评估利用人工智能(AI)方法通过组织病理学数据进行卵巢癌诊断或预后预测的已发表研究,并对其质量进行评价。方法 - 截至2022年12月1日,对5个数据源进行了检索。纳入标准要求研究评估AI在组织病理学图像中进行卵巢癌(包括输卵管-卵巢及腹膜肿瘤)诊断或预后推断的应用。排除综述及非英语文献。使用PROBAST工具对每项纳入模型进行偏倚风险评估。结果 - 共识别出1434篇研究文献,其中36篇符合纳入标准。这些研究报告了62个目标模型,包括35个分类器、14个生存预测模型、7个分割模型和6个回归模型。模型基于1-1375张切片进行开发,涉及1-664例卵巢癌患者。预测结果涵盖总生存期(9/62)、组织学亚型(7/62)、染色强度(6/62)及恶性程度(5/62)等多个指标。早期研究采用基于手工特征的传统机器学习(ML)模型,而近期研究通常使用深度学习(DL)自动学习特征并预测目标结果。所有模型均被评估为整体高偏倚风险或偏倚风险不明确。研究普遍受限于报告不足、样本量小及验证不充分。结论 - 目前相关研究有限,且尚无模型被证明可实际应用于真实场景。针对研究设计中的潜在偏倚与缺陷提出了建议,以促进未来高质量、可重复的研究。关键方面包括更透明全面的报告机制,以及利用交叉验证和外部验证改进的性能评估方法。