Medical multiple-choice question answering (MCQA) is particularly difficult. Questions may describe patient symptoms and ask for the correct diagnosis, which requires domain knowledge and complex reasoning. Standard language modeling pretraining alone is not sufficient to achieve the best results. \citet{jin2020disease} showed that focusing masked language modeling on disease name prediction when using medical encyclopedic paragraphs as input leads to considerable MCQA accuracy improvement. In this work, we show that (1) fine-tuning on generated MCQA dataset outperforms the masked language modeling based objective and (2) correctly masking the cues to the answers is critical for good performance. We release new pretraining datasets and achieve state-of-the-art results on 4 MCQA datasets, notably +5.7\% with base-size model on MedQA-USMLE.
翻译:医学多选题问答(MCQA)是一项极具挑战性的任务。问题可能描述患者症状并要求给出正确诊断,这需要领域知识与复杂推理能力。仅靠标准语言模型预训练不足以获得最佳结果。\citet{jin2020disease}研究表明,当以医学百科全书文本作为输入时,将遮蔽语言模型聚焦于疾病名称预测可显著提升MCQA准确率。在本工作中,我们证明:(1)基于生成的MCQA数据集进行微调优于掩码语言建模目标;(2)正确遮蔽答案线索对模型性能至关重要。我们发布了新的预训练数据集,并在4个MCQA数据集上取得最佳结果,其中在MedQA-USMLE数据集上以基础模型规模实现+5.7%的准确率提升。