Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools, memory, and execution environments. However, this modularity introduces attack surfaces that are largely missed by existing safety evaluations: even when the user request is benign, unsafe influence may reside in skill guidance, local artifacts, or execution-environment files that steer the agent toward unsafe actions. We present SkillSafetyBench, a runnable benchmark for evaluating such skill-mediated safety failures. SkillSafetyBench includes 155 adversarial cases across 47 tasks, 6 risk domains, and 30 safety categories, each evaluated with a case-specific rule-based verifier. Experiments with multiple CLI agents and model backends show that non-user attacks can consistently induce unsafe behavior, with distinct failure patterns across domains, attack methods, and scaffold-model pairings. Our findings suggest that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.


翻译:可复用技能正成为扩展大语言模型智能体的常见接口,它将程序性指导与文件、工具、内存及执行环境的访问权限进行封装。然而,这种模块化机制引入了现有安全性评估往往遗漏的攻击面:即便用户请求本身无害,不安全的影响可能潜藏于技能指导、本地产物或执行环境文件中,从而诱导智能体执行危险操作。我们提出SkillSafetyBench——一个可运行的基准测试,用于评估此类经由技能中介的安全失效问题。该基准包含47项任务、6个风险领域、30个安全类别下的155个对抗性案例,每个案例通过特定规则的验证器进行评估。针对多个CLI智能体及后端模型的实验表明,非用户攻击可持续诱导不安全行为,且不同领域、攻击方法及脚手架-模型组合呈现出差异化的失效模式。研究发现表明,智能体安全性不仅取决于模型层面的对齐,更依赖于智能体如何解读技能、信任工作流上下文以及在可执行环境中的行动决策。

0
下载
关闭预览

相关内容

智能体技能综合综述:分类、技术与应用
专知会员服务
35+阅读 · 5月11日
通用智能体评估的逻辑架构
专知会员服务
22+阅读 · 2月28日
智能体评判者(Agent-as-a-Judge)研究综述
专知会员服务
37+阅读 · 1月9日
智能体安全综述:应用、威胁与防御
专知会员服务
45+阅读 · 2025年10月12日
智能体任务执行安全要求
专知会员服务
19+阅读 · 2025年7月12日
《人工智能安全标准体系(V1.0)》(征求意见稿)
专知会员服务
29+阅读 · 2025年3月23日
《人工智能安全测评白皮书》,99页pdf
专知
36+阅读 · 2022年2月26日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
智能体技能综合综述:分类、技术与应用
专知会员服务
35+阅读 · 5月11日
通用智能体评估的逻辑架构
专知会员服务
22+阅读 · 2月28日
智能体评判者(Agent-as-a-Judge)研究综述
专知会员服务
37+阅读 · 1月9日
智能体安全综述:应用、威胁与防御
专知会员服务
45+阅读 · 2025年10月12日
智能体任务执行安全要求
专知会员服务
19+阅读 · 2025年7月12日
《人工智能安全标准体系(V1.0)》(征求意见稿)
专知会员服务
29+阅读 · 2025年3月23日
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员