We introduce a dataset for evidence/rationale extraction on an extreme multi-label classification task over long medical documents. One such task is Computer-Assisted Coding (CAC) which has improved significantly in recent years, thanks to advances in machine learning technologies. Yet simply predicting a set of final codes for a patient encounter is insufficient as CAC systems are required to provide supporting textual evidence to justify the billing codes. A model able to produce accurate and reliable supporting evidence for each code would be a tremendous benefit. However, a human annotated code evidence corpus is extremely difficult to create because it requires specialized knowledge. In this paper, we introduce MDACE, the first publicly available code evidence dataset, which is built on a subset of the MIMIC-III clinical records. The dataset -- annotated by professional medical coders -- consists of 302 Inpatient charts with 3,934 evidence spans and 52 Profee charts with 5,563 evidence spans. We implemented several evidence extraction methods based on the EffectiveCAN model (Liu et al., 2021) to establish baseline performance on this dataset. MDACE can be used to evaluate code evidence extraction methods for CAC systems, as well as the accuracy and interpretability of deep learning models for multi-label classification. We believe that the release of MDACE will greatly improve the understanding and application of deep learning technologies for medical coding and document classification.
翻译:我们提出一个用于长医学文档极端多标签分类任务中的证据/依据抽取数据集。此类任务之一是计算机辅助编码(CAC),近年来由于机器学习技术的进步取得了显著进展。然而,仅预测患者就诊的最终编码集并不足够,因为CAC系统需要提供支持性文本证据来证明计费编码的合理性。能够为每个编码生成准确可靠支持性证据的模型将带来巨大益处。但由于需要专业知识,人工标注的编码证据语料库极难构建。本文介绍了MDACE——首个公开可用的编码证据数据集,基于MIMIC-III临床记录子集构建。该数据集由专业医学编码员标注,包含302份住院病历(含3,934个证据片段)和52份专业费用病历(含5,563个证据片段)。我们基于EffectiveCAN模型(Liu等人,2021)实现了多种证据抽取方法,以建立该数据集的基线性能。MDACE可用于评估CAC系统的编码证据抽取方法,以及多标签分类深度学习模型的准确性与可解释性。我们相信MDACE的发布将极大推动深度学习技术在医学编码与文档分类领域的理解与应用。