Language models have been shown to perform remarkably well on a wide range of natural language processing tasks. In this paper, we propose LEAP, a novel system that uses language models to perform multi-step logical reasoning and incorporates explicit planning into the inference procedure. Explicit planning enables the system to make more informed reasoning decisions at each step by looking ahead into their future effects. Moreover, we propose a training strategy that safeguards the planning process from being led astray by spurious features. Our full system significantly outperforms other competing methods on multiple standard datasets. When using small T5 models as its core selection and deduction components, our system performs competitively compared to GPT-3 despite having only about 1B parameters (i.e., 175 times smaller than GPT-3). When using GPT-3.5, it significantly outperforms chain-of-thought prompting on the challenging PrOntoQA dataset. We have conducted extensive empirical studies to demonstrate that explicit planning plays a crucial role in the system's performance.
翻译:语言模型在广泛的自然语言处理任务中表现出色。在本文中,我们提出LEAP这一新颖系统,该系统利用语言模型执行多步逻辑推理,并将显式规划融入推理过程。显式规划使系统能够通过前瞻未来影响,在每一步做出更明智的推理决策。此外,我们提出一种训练策略,保护规划过程免受虚假特征的误导。我们的完整系统在多个标准数据集上显著优于其他竞争方法。当使用小型T5模型作为核心选择与推理组件时,尽管仅有约10亿参数(即比GPT-3小175倍),我们的系统仍能与GPT-3竞争。当使用GPT-3.5时,它在具有挑战性的PrOntoQA数据集上显著优于思维链提示方法。我们开展了大量实证研究,证明显式规划在系统性能中发挥关键作用。