Agent skills let LLM agents reuse instructions, resources, tools, and workflows, but they also create a new place for malicious behavior to hide. A skill may look benign in its documentation or code while becoming harmful only when it is invoked with particular user requests, local assets, persistent state, or multi-step tool interactions. This makes purely static vetting brittle. We present Runtime Skill Audit (RSA), a dynamic analysis method that audits skills by asking what the skill-mediated agent actually does under targeted runtime conditions. Instead of testing every skill with the same generic tasks, RSA profiles risk-relevant interfaces, prepares the execution context needed to exercise them, and assigns security labels from the resulting trace evidence. We instantiate RSA on OpenClaw and evaluate it on 100 skills against representative static baselines. RSA achieves 90.0\% accuracy with an 88.0\% true positive rate and an 8.0\% false positive rate, improving accuracy by 13.0 percentage points over the best static baseline. Under self-evolving attacks, static detectors collapse after one or two rounds, while RSA continues to detect 19--20 out of 20 malicious skills across rounds.


翻译:代理技能使大语言模型代理能够复用指令、资源、工具和工作流,但也为恶意行为提供了新的隐藏空间。一项技能可能在文档或代码层面看似无害,但仅在特定用户请求、本地资产、持久化状态或多步骤工具交互被调用时才会产生危害。这使得纯静态审查方法变得脆弱。本文提出运行时代理技能审计(RSA)方法,这是一种动态分析方法,通过询问技能中介代理在目标化运行时条件下实际执行的操作来审计技能。RSA并非用相同通用任务测试所有技能,而是对风险相关接口进行特征分析,构建执行所需上下文,并根据追踪证据分配安全标签。我们在OpenClaw框架上实现RSA,并在100项技能上与代表性静态基线方法进行对比评估。RSA实现了90.0%的准确率,真阳性率为88.0%,假阳性率为8.0%,相较于最优静态基线方法准确率提升13个百分点。面对自演化攻击,静态检测器在一至两轮攻击后即失效,而RSA在各轮攻击中始终能检测出20个恶意技能中的19-20个。

0
下载
关闭预览

相关内容

RSA( RSA (algorithm), Ron Rivest, Adi Shamir and Leonard Adleman )以三位创始人的名字命名。安全性基于有效的素数分解算法的不存在。以安全的非对称加密充当了现代密码体系的骨干。
大型语言模型代理的安全与隐私综述
专知会员服务
30+阅读 · 2024年8月5日
《基于深度学习的实时武器检测系统》
专知会员服务
34+阅读 · 2024年1月22日
深度学习赋能的恶意代码攻防研究进展
专知会员服务
31+阅读 · 2021年4月11日
八个不容错过的 GitHub Copilot 功能!
CSDN
11+阅读 · 2022年9月22日
基于弱监督的视频时序动作检测的介绍
极市平台
30+阅读 · 2019年2月6日
时序异常检测算法概览
论智
29+阅读 · 2018年8月30日
深度学习时代的目标检测算法
炼数成金订阅号
40+阅读 · 2018年3月19日
综述:深度学习时代的目标检测算法
极市平台
27+阅读 · 2018年3月17日
(Python)时序预测的七种方法
云栖社区
10+阅读 · 2018年2月25日
原创 | Attention Modeling for Targeted Sentiment
黑龙江大学自然语言处理实验室
25+阅读 · 2017年11月5日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Arxiv
0+阅读 · 5月14日
VIP会员
最新内容
从采集到决策:美军视角下的战术情报范式重构
专知会员服务
1+阅读 · 今天2:42
《履带式无人地面战车技术发展现状》
专知会员服务
2+阅读 · 今天1:46
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
相关资讯
八个不容错过的 GitHub Copilot 功能!
CSDN
11+阅读 · 2022年9月22日
基于弱监督的视频时序动作检测的介绍
极市平台
30+阅读 · 2019年2月6日
时序异常检测算法概览
论智
29+阅读 · 2018年8月30日
深度学习时代的目标检测算法
炼数成金订阅号
40+阅读 · 2018年3月19日
综述:深度学习时代的目标检测算法
极市平台
27+阅读 · 2018年3月17日
(Python)时序预测的七种方法
云栖社区
10+阅读 · 2018年2月25日
原创 | Attention Modeling for Targeted Sentiment
黑龙江大学自然语言处理实验室
25+阅读 · 2017年11月5日
相关基金
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员