Recent advances in large language models (LLMs) have demonstrated notable progress on many mathematical benchmarks. However, most of these benchmarks only feature problems grounded in junior and senior high school subjects, contain only multiple-choice questions, and are confined to a limited scope of elementary arithmetic operations. To address these issues, this paper introduces an expansive benchmark suite SciBench that aims to systematically examine the reasoning capabilities required for complex scientific problem solving. SciBench contains two carefully curated datasets: an open set featuring a range of collegiate-level scientific problems drawn from mathematics, chemistry, and physics textbooks, and a closed set comprising problems from undergraduate-level exams in computer science and mathematics. Based on the two datasets, we conduct an in-depth benchmark study of two representative LLMs with various prompting strategies. The results reveal that current LLMs fall short of delivering satisfactory performance, with an overall score of merely 35.80%. Furthermore, through a detailed user study, we categorize the errors made by LLMs into ten problem-solving abilities. Our analysis indicates that no single prompting strategy significantly outperforms others and some strategies that demonstrate improvements in certain problem-solving skills result in declines in other skills. We envision that SciBench will catalyze further developments in the reasoning abilities of LLMs, thereby ultimately contributing to scientific research and discovery.
翻译:论文摘要:大语言模型(LLMs)的最新进展在众多数学基准测试中取得了显著进步。然而,这些基准测试中的大多数问题仅局限于初中和高中科目,仅包含选择题,且仅限于初等算术运算的有限范围。为解决这些问题,本文引入了一个扩展的基准测试套件SciBench,旨在系统性地考察复杂科学问题解决所需的推理能力。SciBench包含两个精心整理的数据集:一个开放式数据集,涵盖数学、化学和物理教科书中的一系列大学级科学问题;另一个封闭式数据集,包含计算机科学和数学本科阶段考试中的问题。基于这两个数据集,我们采用多种提示策略对两个代表性大语言模型进行了深入的基准测试研究。结果表明,当前的大语言模型未能达到令人满意的性能,总体得分仅为35.80%。此外,通过详细的用户研究,我们将大语言模型所犯的错误归类为十种问题解决能力。我们的分析表明,没有任何一种提示策略能显著优于其他策略,且某些在特定问题解决技能上表现出改进的策略会导致其他技能下降。我们期望SciBench能进一步推动大语言模型推理能力的发展,从而最终为科学研究和发现做出贡献。