Large Language Models (LLMs) have demonstrated their ability to collaborate effectively with humans in real-world scenarios. However, LLMs are apt to generate hallucinations, i.e., makeup incorrect text and unverified information, which can cause significant damage when deployed for mission-critical tasks. In this paper, we propose a self-check approach based on reverse validation to detect factual errors automatically in a zero-resource fashion. To facilitate future studies and assess different methods, we construct a hallucination detection benchmark, which is generated by ChatGPT and annotated by human annotators. Contrasting previous studies of zero-resource hallucination detection, our method and benchmark concentrate on passage-level detection instead of sentence-level. We empirically evaluate our method and existing zero-resource detection methods on different domains of benchmark to explore the implicit relation between hallucination and training data. Furthermore, we manually analyze some hallucination cases that LLM failed to capture, revealing the shared limitation of zero-resource methods.
翻译:大型语言模型(LLMs)已证明其在现实场景中与人类有效协作的能力。然而,LLMs 容易产生幻觉,即编造不正确的文本和未经核实的信息,在部署于关键任务时可能造成重大损害。本文提出一种基于反向验证的自检方法,以零资源方式自动检测事实错误。为促进后续研究并评估不同方法,我们构建了一个由 ChatGPT 生成并由人工标注的幻觉检测基准。与以往零资源幻觉检测研究不同,我们的方法和基准专注于段落级而非句子级检测。我们通过实验评估了该方法及现有零资源检测方法在不同领域基准上的表现,以探究幻觉与训练数据之间的隐性关系。此外,我们人工分析了 LLMs 未能捕获的部分幻觉案例,揭示了零资源方法的共同局限性。