Large language models (LLMs) have achieved remarkable success across a wide spectrum of tasks; however, they still face limitations in scenarios that demand long-term planning and spatial reasoning. To facilitate this line of research, in this work, we propose a new benchmark, termed $\textbf{P}$ath $\textbf{P}$lanning from $\textbf{N}$atural $\textbf{L}$anguage ($\textbf{PPNL}$). Our benchmark evaluates LLMs' spatial-temporal reasoning by formulating ''path planning'' tasks that require an LLM to navigate to target locations while avoiding obstacles and adhering to constraints. Leveraging this benchmark, we systematically investigate LLMs including GPT-4 via different few-shot prompting methodologies as well as BART and T5 of various sizes via fine-tuning. Our experimental results show the promise of few-shot GPT-4 in spatial reasoning, when it is prompted to reason and act interleavedly, although it still fails to perform long-term temporal reasoning. In contrast, while fine-tuned LLMs achieved impressive results on in-distribution reasoning tasks, they struggled to generalize to larger environments or environments with more obstacles.
翻译:大语言模型(LLMs)在广泛的任务中取得了显著成功,但在需要长期规划与空间推理的场景中仍存在局限性。为促进这一研究方向,本文提出一个新的基准测试,命名为$\textbf{PPNL}$($\textbf{P}$ath $\textbf{P}$lanning from $\textbf{N}$atural $\textbf{L}$anguage,自然语言路径规划)。该基准通过构建"路径规划"任务来评估LLMs的时空推理能力,要求大语言模型在避开障碍物并遵守约束条件的同时导航至目标位置。利用此基准,我们系统研究了采用不同少样本提示方法的GPT-4,以及不同规模的BART和T5微调模型。实验结果表明,当采用交替推理与行动的提示策略时,少样本GPT-4在空间推理方面展现出潜力,但仍无法完成长期时间推理。相比之下,尽管微调后的LLMs在分布内推理任务上取得了令人瞩目的结果,但它们难以泛化至更大规模或包含更多障碍物的环境。