Large Language Models (LLMs) make natural interfaces to factual knowledge, but their usefulness is limited by their tendency to deliver inconsistent answers to semantically equivalent questions. For example, a model might predict both "Anne Redpath passed away in Edinburgh." and "Anne Redpath's life ended in London." In this work, we identify potential causes of inconsistency and evaluate the effectiveness of two mitigation strategies: up-scaling and augmenting the LM with a retrieval corpus. Our results on the LLaMA and Atlas models show that both strategies reduce inconsistency while retrieval augmentation is considerably more efficient. We further consider and disentangle the consistency contributions of different components of Atlas. For all LMs evaluated we find that syntactical form and other evaluation task artifacts impact consistency. Taken together, our results provide a better understanding of the factors affecting the factual consistency of language models.
翻译:大型语言模型(LLMs)构成了事实知识的自然接口,但其对语义等价问题给出不一致回答的倾向限制了其实用性。例如,一个模型可能同时预测"安妮·雷德帕斯在爱丁堡去世"和"安妮·雷德帕斯的生命终结于伦敦"。本研究识别了不一致性的潜在成因,并评估了两种缓解策略的效果:扩大模型规模及通过检索语料库增强语言模型。我们在LLaMA和Atlas模型上的实验表明,两种策略均能减少不一致性,其中检索增强的效率显著更高。我们进一步分析并解构了Atlas不同组件对一致性的贡献。在所评估的所有语言模型中,我们发现句法形式及其他评估任务伪影会影响一致性。综合而言,本研究结果有助于更深入理解影响语言模型事实一致性的因素。