Large Language Models (LLMs) have revolutionized various domains with extensive knowledge and creative capabilities. However, a critical issue with LLMs is their tendency to produce outputs that diverge from factual reality. This phenomenon is particularly concerning in sensitive applications such as medical consultation and legal advice, where accuracy is paramount. In this paper, we introduce the LLM factoscope, a novel Siamese network-based model that leverages the inner states of LLMs for factual detection. Our investigation reveals distinguishable patterns in LLMs' inner states when generating factual versus non-factual content. We demonstrate the LLM factoscope's effectiveness across various architectures, achieving over 96% accuracy in factual detection. Our work opens a new avenue for utilizing LLMs' inner states for factual detection and encourages further exploration into LLMs' inner workings for enhanced reliability and transparency.
翻译:大型语言模型凭借广泛的知识和创造性能力彻底改变了多个领域。然而,LLM的一个关键问题在于其生成与事实不符的输出倾向。这一现象在医疗咨询和法律建议等对准确性要求极高的敏感应用中尤为令人担忧。本文提出了一种基于孪生网络的新型模型——LLM事实探究镜,该模型利用LLM的内部状态进行事实检测。我们的研究发现,LLM在生成事实内容与非事实内容时,其内部状态存在可区分的模式。我们验证了LLM事实探究镜在不同架构上的有效性,在事实检测中实现了超过96%的准确率。这项工作为利用LLM内部状态进行事实检测开辟了新途径,并鼓励进一步探索LLM的内部运作机制,以增强其可靠性和透明度。