In existing memory benchmarks for Large Language Models (LLMs), the evaluated dialogue sessions often lack long-term semantic consistency, and the underlying personas tend to be flat and static. Furthermore, in real-world scenarios, interactions between users and assistants involve more diverse, heterogeneous data streams, such as documents and emails. These shortcomings significantly limit the realism and effectiveness of current evaluations. To address these limitations, we introduce RHELM (Realistic, Heterogeneous, and Evolving Long-term Memory). Driven by meticulously crafted user profiles and a novel LOOP (pLan-rOllout-evOlve-Prune) module, we construct realistic dialogues across diverse interaction scenarios that exhibit dynamic temporal evolution and long-term coherence. Crucially, these dialogues are deeply integrated with heterogeneous external sources synchronized with the user's temporal event trajectory. The resulting benchmark encompasses challenging question-answer pairs spanning seven inquiry types, with each question mapping to at least one of 27 critical memory characteristics that we identify as essential yet underexplored in current research. Comprehensive experiments across full-context models, retrieval-augmented generation (RAG) methods, and representative memory frameworks reveal that contemporary approaches still expose critical weaknesses in complex, real-world settings, particularly in resolving multi-source aggregation and real-world contextual reasoning.
翻译:现有的大语言模型(LLM)长时记忆基准测试中,评估的对话会话往往缺乏长期语义一致性,且其底层人格特征趋于扁平静态。更关键的是,在真实场景中,用户与助手的交互涉及更多样化、更异构的数据流,例如文档和电子邮件。这些缺陷显著制约了当前评估的真实性与有效性。为突破上述局限,我们提出RHELM(Realistic, Heterogeneous, and Evolving Long-term Memory,即现实、异构与演化长时记忆)。该基准依托精心构建的用户画像与新型LOOP(Plan-Rollout-Evolve-Prune,即规划-展开-演化-剪枝)模块,在展现动态时间演化与长期连贯性的多种交互场景中构造出真实对话。尤为重要的是,这些对话与用户时间事件轨迹中同步的异构外部源实现了深度融合。由此生成的基准涵盖七种查询类型的挑战性问答对,每个问题均映射至我们识别出的27个关键记忆特征中的至少一个——这些特征在现有研究中至关重要却尚未充分探索。通过在全上下文模型、检索增强生成(RAG)方法及代表性记忆框架上的综合实验表明,现有方法在复杂的现实场景中仍存在关键弱点,尤其在解决多源聚合与真实世界语境推理时表现突显不足。