We introduce STREET, a unified multi-task and multi-domain natural language reasoning and explanation benchmark. Unlike most existing question-answering (QA) datasets, we expect models to not only answer questions, but also produce step-by-step structured explanations describing how premises in the question are used to produce intermediate conclusions that can prove the correctness of a certain answer. We perform extensive evaluation with popular language models such as few-shot prompting GPT-3 and fine-tuned T5. We find that these models still lag behind human performance when producing such structured reasoning steps. We believe this work will provide a way for the community to better train and test systems on multi-step reasoning and explanations in natural language.
翻译:我们提出STREET,一个统一的多任务、多领域自然语言推理与解释基准。与大多数现有的问答数据集不同,我们期望模型不仅能回答问题,还能生成逐步的结构化解释,描述问题中的前提如何被用于推导中间结论,从而证明某个答案的正确性。我们通过少量样本提示的GPT-3和微调后的T5等主流语言模型进行了广泛评估。研究发现,这些模型在生成此类结构化推理步骤时仍落后于人类表现。我们相信,这项工作将为学界提供一种途径,以更好地训练和测试系统在自然语言中进行多步推理与解释的能力。