Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools, memory, and execution environments. However, this modularity introduces attack surfaces that are largely missed by existing safety evaluations: even when the user request is benign, unsafe influence may reside in skill guidance, local artifacts, or execution-environment files that steer the agent toward unsafe actions. We present SkillSafetyBench, a runnable benchmark for evaluating such skill-mediated safety failures. SkillSafetyBench includes 155 adversarial cases across 47 tasks, 6 risk domains, and 30 safety categories, each evaluated with a case-specific rule-based verifier. Experiments with multiple CLI agents and model backends show that non-user attacks can consistently induce unsafe behavior, with distinct failure patterns across domains, attack methods, and scaffold-model pairings. Our findings suggest that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.
翻译:可复用技能正成为扩展大语言模型智能体的常见接口,它将程序性指导与文件、工具、内存及执行环境的访问权限进行封装。然而,这种模块化机制引入了现有安全性评估往往遗漏的攻击面:即便用户请求本身无害,不安全的影响可能潜藏于技能指导、本地产物或执行环境文件中,从而诱导智能体执行危险操作。我们提出SkillSafetyBench——一个可运行的基准测试,用于评估此类经由技能中介的安全失效问题。该基准包含47项任务、6个风险领域、30个安全类别下的155个对抗性案例,每个案例通过特定规则的验证器进行评估。针对多个CLI智能体及后端模型的实验表明,非用户攻击可持续诱导不安全行为,且不同领域、攻击方法及脚手架-模型组合呈现出差异化的失效模式。研究发现表明,智能体安全性不仅取决于模型层面的对齐,更依赖于智能体如何解读技能、信任工作流上下文以及在可执行环境中的行动决策。