Large Language Models (LLMs) have demonstrated impressive inferential capabilities, with numerous research endeavors devoted to enhancing this capacity through prompting. Despite these efforts, a unified epistemological foundation is still conspicuously absent. Drawing inspiration from Kant's a priori philosophy, we propose the UPAR prompting framework, designed to emulate the structure of human cognition within LLMs. The UPAR framework is delineated into four phases: "Understand", "Plan", "Act", and "Reflect", enabling the extraction of structured information from complex contexts, prior planning of solutions, execution according to plan, and self-reflection. This structure significantly augments the explainability and accuracy of LLM inference, producing a human-understandable and inspectable inferential trajectory. Furthermore, our work offers an epistemological foundation for existing prompting techniques, allowing for a possible systematic integration of these methods. With GPT-4, our approach elevates the accuracy from COT baseline of 22.92% to 58.33% in a challenging subset of GSM8K, and from 67.91% to 75.40% in the causal judgment task. Without using few-shot examples or external tools, UPAR significantly outperforms existing prompting methods on SCIBENCH, a challenging dataset containing collegiate-level mathematics, chemistry, and physics scientific problems.
翻译:大语言模型(LLMs)已展现出令人印象深刻的推理能力,大量研究致力于通过提示增强这一能力。然而,目前仍明显缺乏统一的认识论基础。受康德先验哲学的启发,我们提出UPAR提示框架,旨在模拟人类认知在LLMs中的结构。UPAR框架划分为四个阶段:“理解”、“规划”、“行动”和“反思”,能够从复杂上下文中提取结构化信息、预先规划解决方案、按计划执行以及进行自我反思。该结构显著增强了LLM推理的可解释性和准确性,生成人类可理解且可检查的推理轨迹。此外,我们的工作为现有提示技术提供了认识论基础,使得这些方法可能实现系统性整合。使用GPT-4,我们的方法在GSM8K的一个具有挑战性子集中将准确率从COT基线的22.92%提升至58.33%,在因果判断任务中从67.91%提升至75.40%。在不使用少样本示例或外部工具的情况下,UPAR在SCIBENCH(一个包含大学水平数学、化学和物理科学问题的具有挑战性数据集)上显著优于现有提示方法。