EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records

We present a new text-to-SQL dataset for electronic health records (EHRs). The utterances were collected from 222 hospital staff, including physicians, nurses, insurance review and health records teams, and more. To construct the QA dataset on structured EHR data, we conducted a poll at a university hospital and templatized the responses to create seed questions. Then, we manually linked them to two open-source EHR databases, MIMIC-III and eICU, and included them with various time expressions and held-out unanswerable questions in the dataset, which were all collected from the poll. Our dataset poses a unique set of challenges: the model needs to 1) generate SQL queries that reflect a wide range of needs in the hospital, including simple retrieval and complex operations such as calculating survival rate, 2) understand various time expressions to answer time-sensitive questions in healthcare, and 3) distinguish whether a given question is answerable or unanswerable based on the prediction confidence. We believe our dataset, EHRSQL, could serve as a practical benchmark to develop and assess QA models on structured EHR data and take one step further towards bridging the gap between text-to-SQL research and its real-life deployment in healthcare. EHRSQL is available at https://github.com/glee4810/EHRSQL.

翻译：我们提出了一个面向电子健康记录（EHR）的新型文本到SQL数据集。话语数据来自222名医院工作人员，包括医师、护士、保险审核与健康记录团队等。为了构建基于结构化EHR数据的问答数据集，我们在大学医院开展问卷调查，将回答模板化以生成种子问题。随后，我们手动将这些问题关联至两个开源EHR数据库（MIMIC-III与eICU），并在数据集中纳入不同时间表达式及来自问卷的不可回答的留存问题。本数据集提出了一系列独特挑战：模型需要1）生成反映医院多样化需求的SQL查询语句，包括简单检索及计算存活率等复杂操作；2）理解多种时间表达式以回答医疗保健领域中的时序敏感性问题；3）基于预测置信度区分给定问题是否可回答。我们相信，EHRSQL数据集可作为开发和评估结构化EHR数据问答模型的实用基准，向弥合文本到SQL研究与医疗实际部署之间的鸿沟迈出关键一步。EHRSQL数据集可通过https://github.com/glee4810/EHRSQL获取。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日