Recent advances in large language models (LLMs), such as ChatGPT, have led to highly sophisticated conversation agents. However, these models suffer from "hallucinations," where the model generates false or fabricated information. Addressing this challenge is crucial, particularly with AI-driven platforms being adopted across various sectors. In this paper, we propose a novel method to recognize and flag instances when LLMs perform outside their domain knowledge, and ensuring users receive accurate information. We find that the use of context combined with embedded tags can successfully combat hallucinations within generative language models. To do this, we baseline hallucination frequency in no-context prompt-response pairs using generated URLs as easily-tested indicators of fabricated data. We observed a significant reduction in overall hallucination when context was supplied along with question prompts for tested generative engines. Lastly, we evaluated how placing tags within contexts impacted model responses and were able to eliminate hallucinations in responses with 98.88% effectiveness.
翻译:近年来,大型语言模型(LLMs)如ChatGPT的进展催生了高度复杂的对话代理。然而,这些模型存在“幻觉”问题,即生成虚假或捏造的信息。解决这一挑战至关重要,尤其是在AI驱动平台被各行各业广泛采用的背景下。本文提出了一种新颖方法,用于识别和标记LLMs在其领域知识外运行的情况,确保用户获得准确信息。我们发现,结合上下文与嵌入标签能有效对抗生成式语言模型中的幻觉。为此,我们以生成URL作为易测的捏造数据指标,在无上下文提示-响应对中建立了幻觉频率基准。我们观察到,当向测试生成引擎提供上下文与问题提示时,总体幻觉显著减少。最后,我们评估了上下文中放置标签对模型响应的影响,并以98.88%的有效率消除了响应中的幻觉。