Long-horizon agents rely on memory mechanisms to compress interaction history, but optimizing memory writing faces a distinct credit assignment challenge: a memory update may be rewarded or penalized due to downstream tool failures, noisy observations, or reasoning errors rather than its own contribution. This causally entangled credit can lead agents to discard useful evidence or preserve irrelevant information. We propose HiMPO, a Hindsight-Informed Memory Policy Optimization framework for assigning less-entangled credit to memory-writing actions in long-horizon agents. HiMPO first estimates the local utility of a memory update by comparing the task-relevant information recoverable from the previous and updated memories under the same pre-write state. It then uses hindsight relevance as a bounded retrospective filter that attenuates memory credit when local utility is not supported by the target outcome. The resulting memory-specific advantage is applied only to memory tokens, while trajectory-level rewards optimize the rest of the agent behavior. Across judge-based open-domain tasks and objective compressive-memory QA, HiMPO improves over strong memory-based and RL-based baselines while preserving compressed-context efficiency. Controlled interventions further show that HiMPO reduces blame leakage from tool-induced errors and improves attribution fidelity of memory updates.


翻译:长周期智能体依赖记忆机制压缩交互历史,但优化记忆写入面临独特的信用分配挑战:某次记忆更新可能因下游工具故障、噪声观测或推理错误而获得奖励或惩罚,而非源于其自身贡献。这种因果纠缠的信用会导致智能体丢弃有效证据或保留无关信息。我们提出HiMPO——一种后见知情记忆策略优化框架,旨在为长周期智能体的记忆写入动作分配更少纠缠的信用。HiMPO首先通过比较在相同写入前状态下,从先前记忆与更新记忆中可恢复的任务相关信息,来估计单次记忆更新的局部效用;随后将后见相关性作为有界回溯滤波器,当局部效用未得到目标结果支持时,对记忆信用进行衰减。由此生成的记忆专属优势仅作用于记忆令牌,而轨迹层级奖励则优化智能体的其余行为。在基于评判的开放域任务与客观压缩记忆问答任务中,HiMPO在保持压缩上下文效率的同时,超越了基于强记忆与强强化学习的基线方法。受控干预实验进一步表明,HiMPO能减少由工具诱发错误导致的归责泄漏,并提升记忆更新的归因保真度。

0
下载
关闭预览

相关内容

MMA:多模态记忆智能体
专知会员服务
11+阅读 · 2月19日
下半场思考:基础智能体记忆机制
专知会员服务
22+阅读 · 2月9日
AI智能体时代中的记忆:形式、功能与动态综述
专知会员服务
37+阅读 · 2025年12月16日
用于语音识别的数据增强
AI研习社
24+阅读 · 2019年6月5日
群体智能:新一代人工智能的重要方向
走向智能论坛
12+阅读 · 2017年8月16日
国家自然科学基金
0+阅读 · 2016年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
9+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
9+阅读 · 8月1日
相关基金
国家自然科学基金
0+阅读 · 2016年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员