Purpose: Machine learning models can only be reliably evaluated if training, validation, and test data splits are representative and not affected by the absence of classes of interest. Surgical workflow and instrument recognition tasks are complicated in this manner, because of heavy data imbalances resulting from different lengths of phases and their erratic occurrences. Furthermore, the issue becomes difficult as sub-properties that help define phases, like instrument (co-)occurrence, are usually not considered when defining the split. We argue that such sub-properties must be equally considered. Methods: This work presents a publicly available data visualization tool that enables interactive exploration of dataset splits for surgical phase and instrument recognition. It focuses on the visualization of the occurrence of phases, phase transitions, instruments, and instrument combinations across sets. Particularly, it facilitates the assessment and identification of sub-optimal dataset splits. Results: We performed an analysis of common Cholec80 dataset splits using the proposed application and were able to uncover phase transitions and combinations of instruments that were not represented in one of the sets. Additionally, we outlined possible improvements to the splits. A user study with ten participants demonstrated the ability of participants to solve a selection of data exploration tasks using the proposed application. Conclusion: In highly unbalanced class distributions, special care should be taken with respect to the selection of an appropriate dataset split. Our interactive data visualization tool presents a promising approach for the assessment of dataset splits for surgical phase and instrument recognition. Evaluation results show that it can enhance the development of machine learning models. The application is available at https://cardio-ai.github.io/endovis-ml/ .
翻译:目的:仅当训练集、验证集和测试集的数据划分具有代表性且未受目标类别缺失影响时,机器学习模型才能得到可靠评估。手术工作流与器械识别任务在此方面存在复杂性,原因在于不同阶段的时长差异及其不规则出现导致严重的数据不平衡。此外,由于划分时通常未考虑有助于定义阶段的子属性(如器械共现性),该问题进一步加剧。我们认为必须同等重视此类子属性。方法:本文提出一种公开可用的数据可视化工具,支持对手术阶段与器械识别的数据集划分进行交互式探索。该工具聚焦于可视化各数据集中阶段的出现频率、阶段转换、器械及器械组合的分布情况,尤其有助于评估和识别次优的数据集划分。结果:我们利用所提应用对Cholec80数据集的常见划分进行分析,发现了未被某一数据子集覆盖的阶段转换及器械组合。此外,我们概述了划分的潜在改进方案。十名参与者的人机实验结果表明,用户能够使用该应用完成一系列数据探索任务。结论:在高度不平衡的类别分布中,选择适当的数据集划分需格外谨慎。本文提出的交互式数据可视化工具为评估手术阶段与器械识别中的数据集划分提供了一种有前景的方法。评估结果显示,该工具可增强机器学习模型的开发。应用访问地址:https://cardio-ai.github.io/endovis-ml/