Recent advances in deep learning have led to the development of models approaching the human level of accuracy. However, healthcare remains an area lacking in widespread adoption. The safety-critical nature of healthcare results in a natural reticence to put these black-box deep learning models into practice. This paper explores interpretable methods for a clinical decision support system called sleep staging, an essential step in diagnosing sleep disorders. Clinical sleep staging is an arduous process requiring manual annotation for each 30s of sleep using physiological signals such as electroencephalogram (EEG). Recent work has shown that sleep staging using simple models and an exhaustive set of features can perform nearly as well as deep learning approaches but only for some specific datasets. Moreover, the utility of those features from a clinical standpoint is ambiguous. On the other hand, the proposed framework, NormIntSleep demonstrates exceptional performance across different datasets by representing deep learning embeddings using normalized features. NormIntSleep performs 4.5% better than the exhaustive feature-based approach and 1.5% better than other representation learning approaches. An empirical comparison between the utility of the interpretations of these models highlights the improved alignment with clinical expectations when performance is traded-off slightly. NormIntSleep paired with a clinically meaningful set of features can best balance this trade-off by providing reliable, clinically relevant interpretation with robust performance.
翻译:深度学习的最新进展推动了接近人类水平准确度模型的发展。然而,医疗领域仍是一个广泛应用不足的领域。医疗的安全关键性导致人们自然不愿将这些黑箱深度学习模型付诸实践。本文探讨了临床决策支持系统——睡眠分期的可解释方法,这是诊断睡眠障碍的重要步骤。临床睡眠分期是一项艰巨的任务,需要使用脑电图(EEG)等生理信号对每30秒的睡眠进行手动标注。近期研究表明,使用简单模型和详尽特征集的睡眠分期性能几乎与深度学习方法相当,但仅限于某些特定数据集。此外,从临床角度来看,这些特征的实用性尚不明确。另一方面,提出的框架NormIntSleep通过使用归一化特征表示深度学习嵌入,在不同数据集上展现出卓越性能。与基于详尽特征的方法相比,NormIntSleep性能提升4.5%,比其他表示学习方法提升1.5%。对这些模型解释实用性的实证比较表明,在略微牺牲性能的情况下,其与临床期望的契合度有所提高。当NormIntSleep与一组具有临床意义的特征结合使用时,能通过提供可靠、临床相关的解释和稳健的性能,最好地平衡这一权衡。