Pre-trained Language Models (PLMs) are trained on vast unlabeled data, rich in world knowledge. This fact has sparked the interest of the community in quantifying the amount of factual knowledge present in PLMs, as this explains their performance on downstream tasks, and potentially justifies their use as knowledge bases. In this work, we survey methods and datasets that are used to probe PLMs for factual knowledge. Our contributions are: (1) We propose a categorization scheme for factual probing methods that is based on how their inputs, outputs and the probed PLMs are adapted; (2) We provide an overview of the datasets used for factual probing; (3) We synthesize insights about knowledge retention and prompt optimization in PLMs, analyze obstacles to adopting PLMs as knowledge bases and outline directions for future work.
翻译:预训练语言模型在大量未标注数据上进行训练,蕴含丰富的世界知识。这一事实引起了学界对其所包含事实知识量化研究的兴趣——这既能解释模型在下游任务中的表现,也为将其用作知识库提供潜在依据。本研究系统梳理了用于探针预训练语言模型事实知识的方法与数据集。主要贡献包括:(1)提出基于输入输出调整方式和探针模型适配机制的事实探针方法分类体系;(2)全面概述事实探针研究所用的数据集;(3)综合论述预训练语言模型的知识保留机制与提示优化策略,分析其作为知识库的应用障碍,并指出未来研究方向。