We identify a critical security vulnerability in mainstream Claw personal AI agents: untrusted content encountered during heartbeat-driven background execution can silently pollute agent memory and subsequently influence user-facing behavior without the user's awareness. This vulnerability arises from an architectural design shared across the Claw ecosystem: heartbeat background execution runs in the same session as user-facing conversation, so content ingested from any external source monitored in the background (including email, message channels, news feeds, code repositories, and social platforms) can enter the same memory context used for foreground interaction, often with limited user visibility and without clear source provenance. We formalize this process as an Exposure (E) $\rightarrow$ Memory (M) $\rightarrow$ Behavior (B) pathway: misinformation encountered during heartbeat execution enters the agent's short-term session context, potentially gets written into long-term memory, and later shapes downstream user-facing behavior. We instantiate this pathway in an agent-native social setting using MissClaw, a controlled research replica of Moltbook. We find that (1) social credibility cues, especially perceived consensus, are the dominant driver of short-term behavioral influence, with misleading rates up to 61%; (2) routine memory-saving behavior can promote short-term pollution into durable long-term memory at rates up to 91%, with cross-session behavioral influence reaching 76%; (3) under naturalistic browsing with content dilution and context pruning, pollution still crosses session boundaries. Overall, prompt injection is not required: ordinary social misinformation is sufficient to silently shape agent memory and behavior under heartbeat-driven background execution.


翻译:我们识别出主流Claw个人AI代理中的一个关键安全漏洞:在心跳触发的后台执行过程中,遭遇的不受信任内容能够静默污染代理内存,进而在用户未察觉的情况下影响面向用户的行为。该漏洞源于Claw生态系统共享的架构设计:心跳后台执行与面向用户的对话运行于同一会话中。因此,从后台监控的任何外部源(包括电子邮件、消息通道、新闻推送、代码仓库和社交平台)摄取的内容,均可进入与前台交互相同的内存上下文,且通常用户可见性有限,缺乏明确的来源溯源机制。我们形式化地将此过程定义为曝光(E)→内存(M)→行为(B)路径:心跳执行期间遇到的错误信息进入代理的短期会话上下文,可能被写入长期内存,并在后续影响面向用户的下游行为。我们使用Moltbook的受控研究复刻品MissClaw,在代理原生的社交场景中实例化此路径。研究发现:(1)社交可信度线索,尤其是感知共识,是短期行为影响的主导因素,误导率高达61%;(2)常规的内存保存行为可将短期污染升级为持久性长期内存,比例高达91%,跨会话行为影响达76%;(3)在具有内容稀释与上下文修剪的自然浏览场景下,污染仍能跨越会话边界。总体而言,提示注入并非必要条件:普通社交错误信息足以在心跳驱动的后台执行下静默塑造代理内存与行为。

0
下载
关闭预览

相关内容

⚡ MMClaw: 超轻量级、纯 Python 开发的 AI Agent 内核
专知会员服务
20+阅读 · 2月10日
《使用静态污点分析检测恶意代码》CMU最新30页slides
专知会员服务
22+阅读 · 2023年10月11日
【ICLR2021】神经元注意力蒸馏消除DNN中的后门触发器
专知会员服务
15+阅读 · 2021年1月31日
八个不容错过的 GitHub Copilot 功能!
CSDN
11+阅读 · 2022年9月22日
TF1 到 TF2, 你的在线推理很可能内存爆炸
AINLP
12+阅读 · 2020年6月1日
iOS如何区分App和SDK内部crash
CocoaChina
11+阅读 · 2019年4月17日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
《无人机对海面作战影响评估》
专知会员服务
10+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
5+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
8+阅读 · 7月19日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员