Large language models perform well on many logical reasoning benchmarks, but it remains unclear which core logical skills they truly master. To address this, we introduce LogicSkills, a benchmark that isolates three fundamental logical skills: (i) $\textit{formal symbolization}\unicode{x2014}{}$translating premises into first-order logic; (ii) $\textit{countermodel construction}\unicode{x2014}$showing that an argument is logically invalid by constructing a finite countermodel; and (iii) $\textit{validity assessment}\unicode{x2014}$determining whether a conclusion follows from a set of premises. Items are drawn from the two-variable fragment of first-order logic without identity and are presented in both English and a Carrollian nonce-word language. All instances are solver-verified with Z3 for correctness and non-triviality. Across conventional instruction-tuned LLMs, performance is high on $\textit{validity assessment}$ but substantially lower on $\textit{formal symbolization}$ and $\textit{countermodel construction}$, highlighting that high task-level accuracy can mask weaknesses in core logical skills. In contrast, recent reasoning-tuned models perform strongly across all three tasks, suggesting a more systematic logical skill profile.


翻译:大语言模型在许多逻辑推理基准测试中表现良好,但其真正掌握了哪些核心逻辑技能尚不明确。为此,我们提出了LogicSkills基准,该基准分离出三项基础逻辑技能:(i) $\textit{形式符号化}\unicode{x2014}{}$将前提翻译为一阶逻辑表达式;(ii) $\textit{反模型构建}\unicode{x2014}$通过构建有限反模型来论证逻辑无效性;(iii) $\textit{有效性评估}\unicode{x2014}$判断结论是否从一组前提中得出。测试项选自无等词二元一阶逻辑片段,并以英语和卡罗尔式无意义词语言两种形式呈现。所有实例均通过Z3求解器验证其正确性与非平凡性。在传统的指令微调大语言模型中,$\textit{有效性评估}$任务表现优异,但$\textit{形式符号化}$与$\textit{反模型构建}$任务表现显著较弱,这表明高任务级准确率可能掩盖核心逻辑技能的缺陷。相比之下,近期经过推理调优的模型在所有三项任务中均表现强劲,显示出更为系统化的逻辑技能图谱。

0
下载
关闭预览

相关内容

大语言模型的智能体化推理
专知会员服务
36+阅读 · 1月21日
大语言模型基准综述
专知会员服务
27+阅读 · 2025年8月22日
通过逻辑推理赋能大语言模型:综述
专知会员服务
33+阅读 · 2025年2月24日
大语言模型中的逻辑推理:综述
专知会员服务
49+阅读 · 2025年2月15日
大模型数学推理数据合成相关方法
专知会员服务
36+阅读 · 2025年1月19日
迈向大型推理模型:基于大型语言模型的强化推理综述
专知会员服务
50+阅读 · 2025年1月17日
「大型语言模型推理」综述
专知会员服务
96+阅读 · 2022年12月24日
绝对干货!NLP预训练模型:从transformer到albert
新智元
15+阅读 · 2019年11月10日
自然语言处理(NLP)知识结构总结
AI100
51+阅读 · 2018年8月17日
关系推理:基于表示学习和语义要素
计算机研究与发展
19+阅读 · 2017年8月22日
语料库构建——自然语言理解的基础
计算机研究与发展
11+阅读 · 2017年8月21日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
VIP会员
最新内容
反制无人机:乌克兰提供的五点启示
专知会员服务
4+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
5+阅读 · 9月23日
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
4+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
7+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
9+阅读 · 9月22日
战争不仅需要机器人:人类仍不可或缺
专知会员服务
5+阅读 · 9月21日
《描绘美国防部创新基础设施的未来蓝图》100页
专知会员服务
10+阅读 · 9月21日
相关VIP内容
大语言模型的智能体化推理
专知会员服务
36+阅读 · 1月21日
大语言模型基准综述
专知会员服务
27+阅读 · 2025年8月22日
通过逻辑推理赋能大语言模型:综述
专知会员服务
33+阅读 · 2025年2月24日
大语言模型中的逻辑推理:综述
专知会员服务
49+阅读 · 2025年2月15日
大模型数学推理数据合成相关方法
专知会员服务
36+阅读 · 2025年1月19日
迈向大型推理模型:基于大型语言模型的强化推理综述
专知会员服务
50+阅读 · 2025年1月17日
「大型语言模型推理」综述
专知会员服务
96+阅读 · 2022年12月24日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员