This paper introduces a new testbed CLIFT (Clinical Shift) for the clinical domain Question-answering task. The testbed includes 7.5k high-quality question answering samples to provide a diverse and reliable benchmark. We performed a comprehensive experimental study and evaluated several QA deep-learning models under the proposed testbed. Despite impressive results on the original test set, the performance degrades when applied to new test sets, which shows the distribution shift. Our findings emphasize the need for and the potential for increasing the robustness of clinical domain models under distributional shifts. The testbed offers one way to track progress in that direction. It also highlights the necessity of adopting evaluation metrics that consider robustness to natural distribution shifts. We plan to expand the corpus by adding more samples and model results. The full paper and the updated benchmark are available at github.com/openlifescience-ai/clift
翻译:本文提出了一个面向临床领域问答任务的新型测试平台CLIFT(Clinical Shift)。该测试平台包含7500个高质量问答样本,旨在构建多样化且可靠的基准。我们开展了全面的实验研究,在提出的测试平台上评估了多种深度学习问答模型。尽管模型在原始测试集上取得了显著效果,但应用于新测试集时性能出现下降,这证实了分布偏移的存在。研究结果强调,在分布偏移条件下提升临床领域模型鲁棒性的必要性和潜力。该测试平台为追踪这一方向的进展提供了有效途径,同时凸显了采用考虑自然分布偏移鲁棒性评估指标的重要性。我们计划通过增加样本和模型结果来扩展语料库。完整论文及更新的基准可在 github.com/openlifescience-ai/clift 获取。