Scientific machine reading comprehension (SMRC) aims to understand scientific texts through interactions with humans by given questions. As far as we know, there is only one dataset focused on exploring full-text scientific machine reading comprehension. However, the dataset has ignored the fact that different readers may have different levels of understanding of the text, and only includes single-perspective question-answer pairs, leading to a lack of consideration of different perspectives. To tackle the above problem, we propose a novel multi-perspective SMRC dataset, called SciMRC, which includes perspectives from beginners, students and experts. Our proposed SciMRC is constructed from 741 scientific papers and 6,057 question-answer pairs. Each perspective of beginners, students and experts contains 3,306, 1,800 and 951 QA pairs, respectively. The extensive experiments on SciMRC by utilizing pre-trained models suggest the importance of considering perspectives of SMRC, and demonstrate its challenging nature for machine comprehension.
翻译:科学机器阅读理解(SMRC)旨在通过给定问题与人类互动来理解科学文本。据我们所知,目前仅有一个数据集专注于探索全文本科学机器阅读理解。然而,该数据集忽略了不同读者对文本理解程度可能存在差异这一事实,仅包含单视角问答对,导致缺乏对不同视角的考量。为解决上述问题,我们提出一个新颖的多视角SMRC数据集——SciMRC,其中包含初学者、学生和专家三个视角。所提出的SciMRC基于741篇科学论文构建,包含6,057个问答对。其中,初学者、学生和专家视角分别包含3,306个、1,800个和951个问答对。通过利用预训练模型在SciMRC上进行的大量实验表明,考虑SMRC视角的重要性,并印证了其对机器阅读理解的挑战性。