We present Pre-trained Machine Reader (PMR), a novel method for retrofitting pre-trained masked language models (MLMs) to pre-trained machine reading comprehension (MRC) models without acquiring labeled data. PMR can resolve the discrepancy between model pre-training and downstream fine-tuning of existing MLMs. To build the proposed PMR, we constructed a large volume of general-purpose and high-quality MRC-style training data by using Wikipedia hyperlinks and designed a Wiki Anchor Extraction task to guide the MRC-style pre-training. Apart from its simplicity, PMR effectively solves extraction tasks, such as Extractive Question Answering and Named Entity Recognition. PMR shows tremendous improvements over existing approaches, especially in low-resource scenarios. When applied to the sequence classification task in the MRC formulation, PMR enables the extraction of high-quality rationales to explain the classification process, thereby providing greater prediction explainability. PMR also has the potential to serve as a unified model for tackling various extraction and classification tasks in the MRC formulation.
翻译:我们提出了一种名为预训练机器阅读理解模型(PMR)的新方法,该方法可在无需标注数据的情况下,将预训练的掩码语言模型(MLM)改造为预训练的机器阅读理解(MRC)模型。PMR能够解决现有MLM在模型预训练与下游微调之间存在的差异。为构建所提出的PMR,我们利用维基百科超链接生成了大量通用且高质量的MRC风格训练数据,并设计了维基锚点抽取(Wiki Anchor Extraction)任务来引导MRC风格的预训练。除简洁性外,PMR还能有效解决抽取式问答(Extractive Question Answering)和命名实体识别(Named Entity Recognition)等抽取任务。PMR相较于现有方法展现出显著优势,尤其在低资源场景下表现突出。当在MRC框架下应用于序列分类任务时,PMR能够提取高质量的解释性证据以阐释分类过程,从而提升预测的可解释性。PMR还有望作为统一模型,用于处理MRC框架下的各类抽取与分类任务。