Quantifying uncertainty in automatically generated text is important for letting humans check potential hallucinations and making systems more reliable. Conformal prediction is an attractive framework to provide predictions imbued with statistical guarantees, however, its application to text generation is challenging since any i.i.d. assumptions are not realistic. In this paper, we bridge this gap by leveraging recent results on non-exchangeable conformal prediction, which still ensures bounds on coverage. The result, non-exchangeable conformal nucleus sampling, is a novel extension of the conformal prediction framework to generation based on nearest neighbors. Our method can be used post-hoc for an arbitrary model without extra training and supplies token-level, calibrated prediction sets equipped with statistical guarantees. Experiments in machine translation and language modeling show encouraging results in generation quality. By also producing tighter prediction sets with good coverage, we thus give a more theoretically principled way to perform sampling with conformal guarantees.
翻译:量化自动生成文本中的不确定性对于帮助人类检查潜在的幻觉并使系统更加可靠至关重要。保形预测是一个吸引人的框架,能够提供具有统计保证的预测,然而其在文本生成中的应用具有挑战性,因为任何独立同分布假设都不现实。在本文中,我们通过利用非可交换保形预测的最新研究结果来弥合这一差距,该结果仍能保证覆盖率边界。由此得到的非可交换保形核心采样是保形预测框架基于最近邻的生成的一种新型扩展。我们的方法可以事后应用于任意模型,无需额外训练,并提供带有统计保证的、经过校准的标记级预测集。在机器翻译和语言建模中的实验显示出令人鼓舞的生成质量结果。通过同时生成具有良好覆盖率的更紧凑的预测集,我们因此提供了一种理论上更严谨的方法来执行具有保形保证的采样。