Automatic translation of natural language mathematics into faithful Lean 4 code is hindered by the fundamental dissonance between informal set-theoretic intuition and strict formal type theory. This gap often causes LLMs to hallucinate non-existent library definitions, resulting in code that fails to compile or lacks semantic fidelity. In this work, we investigate the effectiveness of tool-augmented agents for this task through a systematic factorial analysis of three distinct tool categories: Fine-tuned Model Querying (accessing expert drafts), Knowledge Search (retrieving symbol definitions), and Compiler Feedback (verifying code via a Lean REPL). We first benchmark the agent against one-shot baselines, demonstrating large gains in both compilation success and semantic equivalence. We then use the factorial decomposition to quantify the impact of each category, isolating the marginal contribution of each tool type to overall performance.
翻译:自然语言数学到忠实Lean 4代码的自动翻译,受到非正式集合论直觉与严格形式类型论之间根本性不协调的阻碍。这一差距常导致大型语言模型幻觉出不存在的库定义,从而生成无法编译或缺乏语义保真度的代码。在本研究中,我们通过三种不同工具类别的系统因子分析——微调模型查询(获取专家草稿)、知识搜索(检索符号定义)和编译器反馈(通过Lean REPL验证代码)——考察了工具增强型智能体在此任务中的有效性。我们首先将智能体与一次性基线进行基准测试,展示了在编译成功率和语义等价性上的显著提升。接着,我们利用因子分解量化每个类别的影响,并分离出每种工具类型对整体性能的边际贡献。