Intrigued by the claims of emergent reasoning capabilities in LLMs trained on general web corpora, in this paper, we set out to investigate their planning capabilities. We aim to evaluate (1) the effectiveness of LLMs in generating plans autonomously in commonsense planning tasks and (2) the potential of LLMs in LLM-Modulo settings where they act as a source of heuristic guidance for external planners and verifiers. We conduct a systematic study by generating a suite of instances on domains similar to the ones employed in the International Planning Competition and evaluate LLMs in two distinct modes: autonomous and heuristic. Our findings reveal that LLMs' ability to generate executable plans autonomously is rather limited, with the best model (GPT-4) having an average success rate of ~12% across the domains. However, the results in the LLM-Modulo setting show more promise. In the LLM-Modulo setting, we demonstrate that LLM-generated plans can improve the search process for underlying sound planners and additionally show that external verifiers can help provide feedback on the generated plans and back-prompt the LLM for better plan generation.
翻译:受限于在通用网络语料上训练的大型语言模型展现出突发推理能力的说法,本文旨在系统探究其规划能力。我们重点评估:(1) 在常识性规划任务中,LLM自主生成规划方案的有效性;(2) 在LLM-Modulo框架下,LLM作为外部规划器与验证器启发式指导源的潜力。通过构建与国际规划竞赛类似领域的成套实例,我们在自主模式与启发式模式下对LLM进行系统评估。研究发现,LLM自主生成可执行规划方案的能力相当有限,最佳模型GPT-4在各领域平均成功率仅为约12%。但在LLM-Modulo框架下显现出更优前景:我们证实LLM生成的规划可以改善底层可靠规划器的搜索过程,同时表明外部验证器能够对生成方案提供反馈,并通过反向提示机制促进LLM生成更优规划。