We introduce a method to improve the zero-shot reasoning abilities of large language models on general language understanding tasks. Specifically, we build an autonomous agent to instruct the reasoning process of large language models. We show this approach further unleashes the zero-shot reasoning abilities of large language models to more tasks. We study the performance of our method on a wide set of datasets spanning generation, classification, and reasoning. We show that our method generalizes to most tasks and obtains state-of-the-art zero-shot performance on 20 of the 29 datasets that we evaluate. For instance, our method boosts the performance of state-of-the-art large language models by a large margin, including Vicuna-13b (13.3%), Llama-2-70b-chat (23.2%), and GPT-3.5 Turbo (17.0%). Compared to zero-shot chain of thought, our improvement in reasoning is striking, with an average increase of 10.5%. With our method, Llama-2-70b-chat outperforms zero-shot GPT-3.5 Turbo by 10.2%.
翻译:我们提出一种方法,旨在提升大型语言模型在通用语言理解任务上的零样本推理能力。具体而言,我们构建一个自主代理来引导大型语言模型的推理过程。研究表明,该方法能进一步释放大型语言模型在更多任务中的零样本推理潜力。我们在涵盖生成、分类与推理的广泛数据集上评估了所提方法的性能。结果显示,该方法能泛化至大多数任务,并在评估的29个数据集中有20个取得了最先进的零样本性能。例如,该方法显著提升了当前最先进大型语言模型的性能,包括Vicuna-13b(提升13.3%)、Llama-2-70b-chat(提升23.2%)及GPT-3.5 Turbo(提升17.0%)。与零样本思维链相比,我们的推理改进效果显著,平均提升达10.5%。采用该方法后,Llama-2-70b-chat的性能超越零样本GPT-3.5 Turbo达10.2%。