Large language model (LLM) agents increasingly rely on reusable skills i.e. documents describing task-specific procedures. However, this introduces a new attack surface for agents to manage. We study two complementary directions for this threat. First, we evaluate guardian-based defenses: an intermediary LLM agent that acts as a mediator for skill file access (dynamic guardian) or pre-rewrites these files at build time (static guardian). Across three LLM agent families, our guardians cut attack success rate (ASR) by well over half while preserving task utility. Second, we stress test them through attack reframing using four attacks that preserve the malicious instruction but change the phrasing. For non-guardian setup, the reframing pushes the ASR up to 81.4\%, but the dynamic guardian brings it down to 18.6\%, showing that real-time mediation is a robust defense.


翻译:大语言模型(LLM)智能体日益依赖可复用技能(即描述特定任务流程的文档)。然而,这为智能体引入了新的攻击面。我们针对该威胁从两个互补方向展开研究:首先,评估基于守卫的防御机制——一种充当技能文件访问中介的中间层LLM智能体(动态守卫),或在构建阶段预先重写这些文件(静态守卫)。在三个LLM智能体家族中,我们的守卫将攻击成功率(ASR)降低超过一半,同时保持任务效用。其次,通过攻击重构对其进行压力测试,采用四种保留恶意指令但改变措辞的攻击方式。在未部署守卫的场景下,重构使攻击成功率最高达81.4%,而动态守卫将其降至18.6%,表明实时中介是一种稳健的防御策略。

0
下载
关闭预览

相关内容

智能体技能综合综述:分类、技术与应用
专知会员服务
38+阅读 · 5月11日
智能体安全综述:应用、威胁与防御
专知会员服务
46+阅读 · 2025年10月12日
AgentOps综述:分类、挑战与未来方向
专知会员服务
40+阅读 · 2025年8月6日
《大语言模型智能体:方法、应用与挑战综述》
专知会员服务
64+阅读 · 2025年3月28日
面向多智能体博弈对抗的对手建模框架
专知
19+阅读 · 2022年9月28日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
75+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
36+阅读 · 2008年12月31日
VIP会员
最新内容
《异构无人水面艇集群作战自主制导算法》130页
《人工智能能通过美国陆军战争学院吗?》报告
军事域人工智能驱动系统的治理
专知会员服务
4+阅读 · 9月14日
相关基金
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
75+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
36+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员