We present PutnamBench, a new multilingual benchmark for evaluating the ability of neural theorem-provers to solve competition mathematics problems. PutnamBench consists of 1697 hand-constructed formalizations of 640 theorems sourced from the William Lowell Putnam Mathematical Competition, the premier undergraduate-level mathematics competition in North America. All the theorems have formalizations in Lean 4 and Isabelle; a substantial subset also has Coq formalizations. Proving the theorems requires significant problem-solving ability and proficiency in a broad range of topics taught in undergraduate mathematics courses. We use PutnamBench to evaluate several established neural and symbolic theorem-provers. These approaches can only solve a handful of the PutnamBench problems, establishing the benchmark as a difficult open challenge for research on neural theorem-proving. PutnamBench is available at https://github.com/trishullab/PutnamBench.
翻译:本文提出PutnamBench,一个用于评估神经定理证明器解决竞赛数学问题能力的多语言新基准。该基准包含640个源自北美顶级本科生数学竞赛——威廉·洛厄尔·普特南数学竞赛的定理,共手工构建了1697个形式化版本。所有定理均提供Lean 4与Isabelle的形式化表述,其中大部分还包含Coq形式化版本。证明这些定理需要卓越的问题解决能力,并需熟练掌握本科数学课程涵盖的广泛知识领域。我们使用PutnamBench对多个成熟的神经与符号定理证明器进行评估。现有方法仅能解决基准中的少量问题,表明该基准为神经定理证明研究领域确立了一个具有挑战性的开放难题。PutnamBench已发布于https://github.com/trishullab/PutnamBench。