While large pretrained language models (PLMs) demonstrate incredible fluency and performance on many natural language tasks, recent work has shown that well-performing PLMs are very sensitive to what prompts are feed into them. Even when prompts are semantically identical, language models may give very different answers. When considering safe and trustworthy deployments of PLMs we would like their outputs to be consistent under prompts that mean the same thing or convey the same intent. While some work has looked into how state-of-the-art PLMs address this need, they have been limited to only evaluating lexical equality of single- or multi-word answers and do not address consistency of generative text sequences. In order to understand consistency of PLMs under text generation settings, we develop a measure of semantic consistency that allows the comparison of open-ended text outputs. We implement several versions of this consistency metric to evaluate the performance of a number of PLMs on paraphrased versions of questions in the TruthfulQA dataset, we find that our proposed metrics are considerably more consistent than traditional metrics embodying lexical consistency, and also correlate with human evaluation of output consistency to a higher degree.
翻译:尽管大型预训练语言模型(PLMs)在许多自然语言任务中展现出令人难以置信的流畅性和性能,但近期研究表明,性能良好的PLMs对输入的提示(prompts)极为敏感。即使提示在语义上完全相同,语言模型也可能给出截然不同的答案。在考虑PLMs的安全可靠部署时,我们希望它们在表达相同含义或意图的提示下能够输出一致的结果。虽然已有一些研究探讨了最先进的PLMs如何满足这一需求,但这些工作仅限于评估单词或多词答案的词汇相等性,并未解决生成式文本序列的一致性。为了理解PLMs在文本生成情境下的一致性,我们开发了一种语义一致性度量方法,允许对开放式文本输出进行比较。我们实现了该一致性度量的多个版本,以评估多个PLMs在TruthfulQA数据集问题改写版本上的性能。研究发现,我们提出的度量方法比体现词汇一致性的传统度量方法具有显著更高的一致性,并且与人工评估的输出一致性结果的相关性也更强。