We identify a critical security vulnerability in mainstream Claw personal AI agents: untrusted content encountered during heartbeat-driven background execution can silently pollute agent memory and subsequently influence user-facing behavior without the user's awareness. This vulnerability arises from an architectural design shared across the Claw ecosystem: heartbeat background execution runs in the same session as user-facing conversation, so content ingested from any external source monitored in the background (including email, message channels, news feeds, code repositories, and social platforms) can enter the same memory context used for foreground interaction, often with limited user visibility and without clear source provenance. We formalize this process as an Exposure (E) $\rightarrow$ Memory (M) $\rightarrow$ Behavior (B) pathway: misinformation encountered during heartbeat execution enters the agent's short-term session context, potentially gets written into long-term memory, and later shapes downstream user-facing behavior. We instantiate this pathway in an agent-native social setting using MissClaw, a controlled research replica of Moltbook. We find that (1) social credibility cues, especially perceived consensus, are the dominant driver of short-term behavioral influence, with misleading rates up to 61%; (2) routine memory-saving behavior can promote short-term pollution into durable long-term memory at rates up to 91%, with cross-session behavioral influence reaching 76%; (3) under naturalistic browsing with content dilution and context pruning, pollution still crosses session boundaries. Overall, prompt injection is not required: ordinary social misinformation is sufficient to silently shape agent memory and behavior under heartbeat-driven background execution.
翻译:我们识别了主流Claw个人AI代理中的一个关键安全漏洞:在心跳驱动的后台执行过程中遇到的不可信内容会静默污染代理内存,并在用户不知情的情况下影响面向用户的行为。该漏洞源于Claw生态系统中共享的架构设计:心跳后台执行与面向用户的对话在同一个会话中运行,因此从后台监控的任何外部来源(包括电子邮件、消息频道、新闻推送、代码仓库和社交平台)摄取的内容都可能进入用于前台交互的同一内存上下文,且通常用户可见性有限,来源追溯不明确。我们将此过程形式化为暴露(E)→ 内存(M)→ 行为(B)通路:心跳执行期间遇到的错误信息进入代理的短期会话上下文,可能被写入长期内存,并随后影响下游面向用户的行为。我们使用Moltbook的受控研究副本MissClaw,在代理原生社交环境中实例化此通路。我们发现:(1)社交可信度线索,尤其是感知共识,是短期行为影响的主要驱动因素,误导率高达61%;(2)常规内存保存行为会将短期污染提升为持久长期内存,污染率高达91%,跨会话行为影响达到76%;(3)在伴有内容稀释和上下文修剪的自然浏览场景下,污染仍会跨越会话边界。总体而言,提示注入并非必需:在心跳驱动的后台执行下,普通的社会错误信息足以静默塑造代理内存和行为。