Our Visual Analytics (VA) tool ScrutinAI supports human analysts to investigate interactively model performanceand data sets. Model performance depends on labeling quality to a large extent. In particular in medical settings, generation of high quality labels requires in depth expert knowledge and is very costly. Often, data sets are labeled by collecting opinions of groups of experts. We use our VA tool to analyse the influence of label variations between different experts on the model performance. ScrutinAI facilitates to perform a root cause analysis that distinguishes weaknesses of deep neural network (DNN) models caused by varying or missing labeling quality from true weaknesses. We scrutinize the overall detection of intracranial hemorrhages and the more subtle differentiation between subtypes in a publicly available data set.
翻译:我们的可视化分析(VA)工具ScrutinAI支持人类分析人员交互式地研究模型性能与数据集。模型性能在很大程度上依赖于标注质量。特别是在医疗场景中,生成高质量标签需要深入的专家知识且成本极高。通常,数据集标注需要收集专家组的意见。我们利用VA工具分析不同专家间标注差异对模型性能的影响。ScrutinAI有助于开展根本原因分析,将深度神经网络(DNN)模型因标注质量差异或缺失导致的缺陷与真正的性能弱点区分开来。我们在一公开数据集中对颅内出血的总体检测及亚型间的细微区分进行了深入分析。