Unsupervised question answering is a promising yet challenging task, which alleviates the burden of building large-scale annotated data in a new domain. It motivates us to study the unsupervised multiple-choice question answering (MCQA) problem. In this paper, we propose a novel framework designed to generate synthetic MCQA data barely based on contexts from the universal domain without relying on any form of manual annotation. Possible answers are extracted and used to produce related questions, then we leverage both named entities (NE) and knowledge graphs to discover plausible distractors to form complete synthetic samples. Experiments on multiple MCQA datasets demonstrate the effectiveness of our method.
翻译:无监督问答是一项具有前景但具有挑战性的任务,它减轻了在新领域构建大规模标注数据的负担。这促使我们研究无监督多项选择题问答(MCQA)问题。本文提出了一种新颖的框架,该框架仅基于通用领域的上下文生成合成MCQA数据,无需依赖任何形式的人工标注。我们提取可能的答案并用于生成相关问题,随后利用命名实体(NE)和知识图谱发现合理的干扰项,以形成完整的合成样本。在多个MCQA数据集上的实验证明了我们方法的有效性。