To explain the predicted answers and evaluate the reasoning abilities of models, several studies have utilized underlying reasoning (UR) tasks in multi-hop question answering (QA) datasets. However, it remains an open question as to how effective UR tasks are for the QA task when training models on both tasks in an end-to-end manner. In this study, we address this question by analyzing the effectiveness of UR tasks (including both sentence-level and entity-level tasks) in three aspects: (1) QA performance, (2) reasoning shortcuts, and (3) robustness. While the previous models have not been explicitly trained on an entity-level reasoning prediction task, we build a multi-task model that performs three tasks together: sentence-level supporting facts prediction, entity-level reasoning prediction, and answer prediction. Experimental results on 2WikiMultiHopQA and HotpotQA-small datasets reveal that (1) UR tasks can improve QA performance. Using four debiased datasets that are newly created, we demonstrate that (2) UR tasks are helpful in preventing reasoning shortcuts in the multi-hop QA task. However, we find that (3) UR tasks do not contribute to improving the robustness of the model on adversarial questions, such as sub-questions and inverted questions. We encourage future studies to investigate the effectiveness of entity-level reasoning in the form of natural language questions (e.g., sub-question forms).
翻译:为解释预测答案并评估模型的推理能力,多项研究在多跳问答数据集中采用了潜在推理(UR)任务。然而,当以端到端方式同时训练这两个任务时,UR任务对问答任务的有效性仍是一个开放性问题。本研究从三个方面分析UR任务(包括句子级和实体级任务)的有效性:(1)问答性能,(2)推理捷径,以及(3)鲁棒性。尽管先前模型未明确训练过实体级推理预测任务,我们构建了一个多任务模型,该模型同时执行三项任务:句子级支持事实预测、实体级推理预测和答案预测。在2WikiMultiHopQA和HotpotQA-small数据集上的实验结果表明:(1)UR任务可提升问答性能。通过使用新创建的四个去偏数据集,我们证明(2)UR任务有助于防止多跳问答任务中的推理捷径。然而,我们发现(3)UR任务并未提升模型对对抗性问题(如子问题与反转问题)的鲁棒性。我们鼓励未来研究以自然语言问题形式(例如子问题形式)探究实体级推理的有效性。