The widespread deployment of LLM-based agents is likely to introduce a critical privacy threat: malicious agents that proactively engage others in multi-turn interactions to extract sensitive information. However, the evolving nature of such dynamic dialogues makes it challenging to anticipate emerging vulnerabilities and design effective defenses. To tackle this problem, we present a search-based framework that alternates between improving attack and defense strategies through the simulation of privacy-critical agent interactions. Specifically, we employ LLMs as optimizers to analyze simulation trajectories and iteratively propose new agent instructions. To explore the strategy space more efficiently, we further utilize parallel search with multiple threads and cross-thread propagation. Through this process, we find that attack strategies escalate from direct requests to sophisticated tactics, such as impersonation and consent forgery, while defenses evolve from simple rule-based constraints to robust identity-verification state machines. The discovered attacks and defenses generalize across diverse scenarios and backbone models, providing useful insights for developing privacy-aware agents.
翻译:基于大语言模型的智能代理的广泛部署可能引发一项关键隐私威胁:恶意代理主动在多方交互中通过多轮对话套取敏感信息。然而,此类动态对话的演变特性使得预测新兴漏洞并设计有效防御策略极具挑战性。针对该问题,我们提出一种基于搜索的框架,通过在隐私关键型代理交互仿真中交替优化攻击与防御策略。具体而言,我们采用大语言模型作为优化器,分析仿真轨迹并迭代提出新代理指令。为更高效探索策略空间,我们进一步利用多线程并行搜索与跨线程传播机制。通过该流程发现:攻击策略从直接请求升级为伪装身份、伪造授权等复杂手段,而防御策略则从简单规则约束发展为具备身份验证功能的状态机。已识别的攻击与防御方法可泛化至不同场景与骨干模型,为开发具备隐私意识的智能代理提供重要参考。