We identify a critical security vulnerability in mainstream Claw personal AI agents: untrusted content encountered during heartbeat-driven background execution can silently pollute agent memory and subsequently influence user-facing behavior without the user's awareness. This vulnerability arises from an architectural design shared across the Claw ecosystem: heartbeat background execution runs in the same session as user-facing conversation, so content ingested from any external source monitored in the background (including email, message channels, news feeds, code repositories, and social platforms) can enter the same memory context used for foreground interaction, often with limited user visibility and without clear source provenance. We formalize this process as an Exposure (E) $\rightarrow$ Memory (M) $\rightarrow$ Behavior (B) pathway: misinformation encountered during heartbeat execution enters the agent's short-term session context, potentially gets written into long-term memory, and later shapes downstream user-facing behavior. We instantiate this pathway in an agent-native social setting using MissClaw, a controlled research replica of Moltbook. We find that (1) social credibility cues, especially perceived consensus, are the dominant driver of short-term behavioral influence, with misleading rates up to 61%; (2) routine memory-saving behavior can promote short-term pollution into durable long-term memory at rates up to 91%, with cross-session behavioral influence reaching 76%; (3) under naturalistic browsing with content dilution and context pruning, pollution still crosses session boundaries. Overall, prompt injection is not required: ordinary social misinformation is sufficient to silently shape agent memory and behavior under heartbeat-driven background execution.
翻译:我们识别出主流Claw个人AI代理中的一个关键安全漏洞:在心跳触发的后台执行过程中,遭遇的不受信任内容能够静默污染代理内存,进而在用户未察觉的情况下影响面向用户的行为。该漏洞源于Claw生态系统共享的架构设计:心跳后台执行与面向用户的对话运行于同一会话中。因此,从后台监控的任何外部源(包括电子邮件、消息通道、新闻推送、代码仓库和社交平台)摄取的内容,均可进入与前台交互相同的内存上下文,且通常用户可见性有限,缺乏明确的来源溯源机制。我们形式化地将此过程定义为曝光(E)→内存(M)→行为(B)路径:心跳执行期间遇到的错误信息进入代理的短期会话上下文,可能被写入长期内存,并在后续影响面向用户的下游行为。我们使用Moltbook的受控研究复刻品MissClaw,在代理原生的社交场景中实例化此路径。研究发现:(1)社交可信度线索,尤其是感知共识,是短期行为影响的主导因素,误导率高达61%;(2)常规的内存保存行为可将短期污染升级为持久性长期内存,比例高达91%,跨会话行为影响达76%;(3)在具有内容稀释与上下文修剪的自然浏览场景下,污染仍能跨越会话边界。总体而言,提示注入并非必要条件:普通社交错误信息足以在心跳驱动的后台执行下静默塑造代理内存与行为。