While Large Language Models (LLMs) can serve as agents to simulate human behaviors (i.e., role-playing agents), we emphasize the importance of point-in-time role-playing. This situates characters at specific moments in the narrative progression for three main reasons: (i) enhancing users' narrative immersion, (ii) avoiding spoilers, and (iii) fostering engagement in fandom role-playing. To accurately represent characters at specific time points, agents must avoid character hallucination, where they display knowledge that contradicts their characters' identities and historical timelines. We introduce TimeChara, a new benchmark designed to evaluate point-in-time character hallucination in role-playing LLMs. Comprising 10,895 instances generated through an automated pipeline, this benchmark reveals significant hallucination issues in current state-of-the-art LLMs (e.g., GPT-4o). To counter this challenge, we propose Narrative-Experts, a method that decomposes the reasoning steps and utilizes narrative experts to reduce point-in-time character hallucinations effectively. Still, our findings with TimeChara highlight the ongoing challenges of point-in-time character hallucination, calling for further study.
翻译:尽管大语言模型(LLMs)可作为模拟人类行为的智能体(即角色扮演智能体),我们强调时点角色扮演的重要性。该方法将角色置于叙事进程的特定时刻,主要基于三个原因:(i)增强用户的叙事沉浸感,(ii)避免剧透,以及(iii)促进粉丝社群的角色扮演参与度。为准确呈现特定时点的角色,智能体必须避免角色幻觉——即表现出与其角色身份及历史时间线相矛盾的知识。我们提出TimeChara,这是一个专为评估角色扮演大语言模型中时点角色幻觉而设计的新基准。该基准包含通过自动化流程生成的10,895个测试实例,揭示了当前最先进大语言模型(如GPT-4o)中普遍存在的显著幻觉问题。为应对这一挑战,我们提出Narrative-Experts方法,该方法通过分解推理步骤并利用叙事专家来有效降低时点角色幻觉。尽管如此,我们基于TimeChara的研究结果仍凸显了时点角色幻觉问题的持续挑战,亟待进一步深入研究。