This report describes a test of the large language model GPT-4 with the Wolfram Alpha and the Code Interpreter plug-ins on 105 original problems in science and math, at the high school and college levels, carried out in June-August 2023. Our tests suggest that the plug-ins significantly enhance GPT's ability to solve these problems. Having said that, there are still often "interface" failures; that is, GPT often has trouble formulating problems in a way that elicits useful answers from the plug-ins. Fixing these interface failures seems like a central challenge in making GPT a reliable tool for college-level calculation problems.
翻译:本报告描述了2023年6月至8月期间,针对大型语言模型GPT-4在105道高中及大学水平的科学与数学原创问题中,结合Wolfram Alpha插件与代码解释器插件的测试结果。我们的测试表明,这些插件显著提升了GPT解决此类问题的能力。但与此同时也存在持续的"接口"故障:即GPT在将问题转化为能有效调用插件提供有用答案的表达方式时经常存在困难。解决这类接口故障,似乎是使GPT成为大学级别计算题可靠工具的核心挑战。