When people think of everyday things like an egg, they typically have a mental image associated with it. This allows them to correctly judge, for example, that "the yolk surrounds the shell" is a false statement. Do language models similarly have a coherent picture of such everyday things? To investigate this, we propose a benchmark dataset consisting of 100 everyday things, their parts, and the relationships between these parts, expressed as 11,720 "X relation Y?" true/false questions. Using these questions as probes, we observe that state-of-the-art pre-trained language models (LMs) like GPT-3 and Macaw have fragments of knowledge about these everyday things, but do not have fully coherent "parts mental models" (54-59% accurate, 19-43% conditional constraint violation). We propose an extension where we add a constraint satisfaction layer on top of the LM's raw predictions to apply commonsense constraints. As well as removing inconsistencies, we find that this also significantly improves accuracy (by 16-20%), suggesting how the incoherence of the LM's pictures of everyday things can be significantly reduced.
翻译:人们在思考鸡蛋等日常事物时,通常会形成与之关联的心理图像。这种能力使他们能够正确判断诸如"蛋黄包裹蛋壳"这类陈述的真伪。语言模型是否同样具备关于此类日常事物的连贯图景?为探究该问题,我们构建了一个基准数据集,包含100个日常事物及其组成部分,以及这些组成部分之间的关联关系,具体表现为11,720个"X 关系 Y?"的真假判断题。通过将这些判断题作为探测工具,我们观察到GPT-3和Macaw等最先进的预训练语言模型(LM)虽掌握了关于这些日常事物的碎片化知识,但尚未形成完全连贯的"部件心理模型"(准确率54-59%,条件约束违反率19-43%)。为此,我们提出扩展方案:在语言模型原始预测结果之上增加约束满足层,用以施加常识性约束。实验表明,该方法不仅消除了不一致性,还将准确率大幅提升16-20%,这证明语言模型关于日常事物的图景不连贯性可被显著改善。