Recent advancements in large language models (LLMs) have showcased significant improvements in mathematics. However, traditional math benchmarks like GSM8k offer a unidimensional perspective, falling short in providing a holistic assessment of the LLMs' math capabilities. To address this gap, we introduce MathBench, a new benchmark that rigorously assesses the mathematical capabilities of large language models. MathBench spans a wide range of mathematical disciplines, offering a detailed evaluation of both theoretical understanding and practical problem-solving skills. The benchmark progresses through five distinct stages, from basic arithmetic to college mathematics, and is structured to evaluate models at various depths of knowledge. Each stage includes theoretical questions and application problems, allowing us to measure a model's mathematical proficiency and its ability to apply concepts in practical scenarios. MathBench aims to enhance the evaluation of LLMs' mathematical abilities, providing a nuanced view of their knowledge understanding levels and problem solving skills in a bilingual context. The project is released at https://github.com/open-compass/MathBench .
翻译:近年来,大语言模型在数学领域展现出显著进步。然而,传统数学基准(如GSM8k)仅提供单一维度的视角,难以全面评估大语言模型的数学能力。为解决这一局限,我们提出MathBench——一项用于严格评估大语言模型数学能力的新基准。MathBench涵盖广泛的数学学科,提供对理论理解与实际问题解决能力的细致评估。该基准从基础算术到高等数学共分为五个递进阶段,旨在评估模型在不同知识深度下的表现。每个阶段包含理论题与应用题,从而衡量模型的数学熟练度及其在实践场景中应用概念的能力。MathBench致力于提升对大语言模型数学能力的评估,在双语语境下揭示其知识理解层次与问题解决技能的细微差异。项目已发布于https://github.com/open-compass/MathBench。