Agent skills are increasingly used to extend LLM agents with task-specific instructions, executable scripts, and auxiliary resources. While improving reusability, this modular design also introduces a new supply-chain attack surface: a malicious or compromised skill may be repeatedly loaded as trusted guidance and steer an agent's tool use during downstream execution. Existing skill-based prompt-injection attacks are mostly manual and brittle, as explicit malicious instructions are often rejected or ignored when poorly aligned with the original skill workflow. We propose SkillJect, the first automated framework for generating effective poisoned skills against skill-enabled agent systems. SkillJect decomposes the attack into two coordinated channels. In the artifact channel, it hides the malicious payload in an auxiliary helper script. In the instruction channel, it rewrites SKILL.md using a front-loaded inducement strategy, placing injected content at the beginning and framing the helper script as a mandatory prerequisite or first step. The instruction explicitly references the helper-script path and provides an executable command, making the helper appear to be a legitimate initialization step before normal operations. SkillJect further adopts a closed-loop multi-agent process to improve attack performance. An Attack Agent generates poisoned skills, a Victim Agent executes downstream tasks with them, and an Evaluate Agent inspects execution traces to determine whether the hidden payload is executed. The Attack Agent then uses this feedback to diagnose failures and rewrite SKILL.md, while keeping the payload fixed. Experiments across platforms, backend LLMs, and attack categories show that SkillJect substantially outperforms naive direct injection and prior manual attacks, revealing poisoned skills as a persistent attack vector in reusable skill ecosystems.


翻译:智能体技能越来越多地用于扩展大语言模型智能体,为其提供任务特定指令、可执行脚本和辅助资源。虽然这种模块化设计提升了可复用性,但也引入了一种新的供应链攻击面:恶意或受损的技能可能被反复加载为可信指导,并在下游执行过程中操控智能体的工具使用。现有的基于技能的提示注入攻击大多为手动且脆弱,因为当显式恶意指令与原技能工作流不一致时,常会被拒绝或忽略。我们提出SkillJect——首个针对技能型智能体系统自动生成有效中毒技能的框架。SkillJect将攻击分解为两个协调通道:在工件通道中,它将恶意载荷隐藏在辅助脚本内;在指令通道中,它采用前置诱导策略重写SKILL.md,将注入内容置于开头,并将辅助脚本框架化为强制前置条件或第一步。指令显式引用辅助脚本路径并提供可执行命令,使该辅助脚本在正常操作前呈现为合法的初始化步骤。SkillJect进一步采用闭环多智能体流程提升攻击性能:攻击智能体生成中毒技能,受害者智能体使用中毒技能执行下游任务,评估智能体检查执行轨迹以判断隐藏载荷是否被执行。攻击智能体随后利用该反馈诊断失败原因并重写SKILL.md,同时保持载荷固定。跨平台、后端大语言模型和攻击类别的实验表明,SkillJect显著优于朴素直接注入和先前手动攻击,揭示了中毒技能作为可复用技能生态系统中持续性攻击向量的威胁。

0
下载
关闭预览

相关内容

智能体技能综合综述:分类、技术与应用
专知会员服务
35+阅读 · 5月11日
伯克利最新《智能体 AI (Agentic AI)》课程
专知会员服务
49+阅读 · 3月1日
《针对指挥控制强化学习智能体的对抗攻击》
专知会员服务
32+阅读 · 2月5日
智能体工程(Agent Engineering)
专知会员服务
39+阅读 · 2025年12月31日
AI智能体编程:技术、挑战与机遇综述
专知会员服务
49+阅读 · 2025年8月18日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Arxiv
0+阅读 · 6月16日
Arxiv
0+阅读 · 6月15日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
6+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关VIP内容
相关资讯
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员