In existing memory benchmarks for Large Language Models (LLMs), the evaluated dialogue sessions often lack long-term semantic consistency, and the underlying personas tend to be flat and static. Furthermore, in real-world scenarios, interactions between users and assistants involve more diverse, heterogeneous data streams, such as documents and emails. These shortcomings significantly limit the realism and effectiveness of current evaluations. To address these limitations, we introduce RHELM (Realistic, Heterogeneous, and Evolving Long-term Memory). Driven by meticulously crafted user profiles and a novel LOOP (pLan-rOllout-evOlve-Prune) module, we construct realistic dialogues across diverse interaction scenarios that exhibit dynamic temporal evolution and long-term coherence. Crucially, these dialogues are deeply integrated with heterogeneous external sources synchronized with the user's temporal event trajectory. The resulting benchmark encompasses challenging question-answer pairs spanning seven inquiry types, with each question mapping to at least one of 27 critical memory characteristics that we identify as essential yet underexplored in current research. Comprehensive experiments across full-context models, retrieval-augmented generation (RAG) methods, and representative memory frameworks reveal that contemporary approaches still expose critical weaknesses in complex, real-world settings, particularly in resolving multi-source aggregation and real-world contextual reasoning.
翻译:现有面向大语言模型(LLMs)的记忆基准测试中,评估的对话会话往往缺乏长期语义一致性,且底层人物设定趋于扁平静态。此外,在真实场景中,用户与助手的交互涉及更多样化的异构数据流(如文档和邮件)。这些缺陷严重限制了当前评估的现实性与有效性。为应对上述局限,我们提出RHELM(现实、异构且演化的长时记忆)基准。借助精心设计的用户画像与新型LOOP(规划-推演-演化-修剪)模块,我们构建了跨多种交互场景的现实对话,这些对话展现出动态时序演化与长期连贯性。关键之处在于,这些对话与用户时间事件轨迹同步的异构外部源实现了深度整合。最终基准包含覆盖七种查询类型的挑战性问答对,每个问题对应至少一项我们识别出的27项关键记忆特征——这些特征虽至关重要却在当前研究中未获充分探索。基于全上下文模型、检索增强生成(RAG)方法与代表性记忆框架的综合实验揭示:当代方法在复杂的真实场景中仍暴露出关键缺陷,尤其在解决多源聚合与真实世界上下文推理任务时。