Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools, memory, and execution environments. However, this modularity introduces attack surfaces that are largely missed by existing safety evaluations: even when the user request is benign, unsafe influence may reside in skill guidance, local artifacts, or execution-environment files that steer the agent toward unsafe actions. We present SkillSafetyBench, a runnable benchmark for evaluating such skill-mediated safety failures. SkillSafetyBench includes 155 adversarial cases across 47 tasks, 6 risk domains, and 30 safety categories, each evaluated with a case-specific rule-based verifier. Experiments with multiple CLI agents and model backends show that non-user attacks can consistently induce unsafe behavior, with distinct failure patterns across domains, attack methods, and scaffold-model pairings. Our findings suggest that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.


翻译:可复用技能正成为扩展大语言模型智能体的常见接口,它将程序性指导与文件、工具、内存及执行环境的访问权限进行封装。然而,这种模块化机制引入了现有安全性评估往往遗漏的攻击面:即便用户请求本身无害,不安全的影响可能潜藏于技能指导、本地产物或执行环境文件中,从而诱导智能体执行危险操作。我们提出SkillSafetyBench——一个可运行的基准测试,用于评估此类经由技能中介的安全失效问题。该基准包含47项任务、6个风险领域、30个安全类别下的155个对抗性案例,每个案例通过特定规则的验证器进行评估。针对多个CLI智能体及后端模型的实验表明,非用户攻击可持续诱导不安全行为,且不同领域、攻击方法及脚手架-模型组合呈现出差异化的失效模式。研究发现表明,智能体安全性不仅取决于模型层面的对齐,更依赖于智能体如何解读技能、信任工作流上下文以及在可执行环境中的行动决策。

0
下载
关闭预览

相关内容

智能体技能综合综述:分类、技术与应用
专知会员服务
38+阅读 · 5月11日
通用智能体评估的逻辑架构
专知会员服务
23+阅读 · 2月28日
智能体评判者(Agent-as-a-Judge)研究综述
专知会员服务
38+阅读 · 1月9日
智能体任务执行安全要求
专知会员服务
20+阅读 · 2025年7月12日
《人工智能安全标准体系(V1.0)》(征求意见稿)
专知会员服务
29+阅读 · 2025年3月23日
《人工智能安全测评白皮书》,99页pdf
专知
36+阅读 · 2022年2月26日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
《异构无人水面艇集群作战自主制导算法》130页
《人工智能能通过美国陆军战争学院吗?》报告
军事域人工智能驱动系统的治理
专知会员服务
4+阅读 · 9月14日
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员