Reading comprehension continues to be a crucial research focus in the NLP community. Recent advances in Machine Reading Comprehension (MRC) have mostly centered on literal comprehension, referring to the surface-level understanding of content. In this work, we focus on the next level - interpretive comprehension, with a particular emphasis on inferring the themes of a narrative text. We introduce the first dataset specifically designed for interpretive comprehension of educational narratives, providing corresponding well-edited theme texts. The dataset spans a variety of genres and cultural origins and includes human-annotated theme keywords with varying levels of granularity. We further formulate NLP tasks under different abstractions of interpretive comprehension toward the main idea of a story. After conducting extensive experiments with state-of-the-art methods, we found the task to be both challenging and significant for NLP research. The dataset and source code have been made publicly available to the research community at https://github.com/RiTUAL-UH/EduStory.
翻译:阅读理解依然是自然语言处理学界的重要研究课题。近年来,机器阅读理解领域的研究进展主要集中于字面理解层面,即对内容的表层理解。本研究聚焦于更高层次的阐释性理解,特别关注从叙事文本中推断主题的能力。我们首次构建了专为教育叙事文本阐释性理解设计的数据集,并提供了经过精细编辑的主题文本。该数据集涵盖多种文体类型和文化起源,包含由人工标注、粒度可调的主题关键词。我们进一步针对故事主旨的阐释性理解,在抽象程度不同的维度上定义了相应的自然语言处理任务。通过采用当前最先进方法进行大量实验,我们发现该任务对自然语言处理研究既具挑战性又具重要意义。数据集与源代码已通过https://github.com/RiTUAL-UH/EduStory 向研究社区公开。