Skills are increasingly used to extend LLM agents by packaging prompts, code, and configurations into reusable modules. As public registries and marketplaces expand, they form an emerging agentic supply chain, but also introduce a new attack surface for malicious skills. Detecting malicious skills is challenging because relevant evidence is often distributed across heterogeneous artifacts and must be reasoned in context. Existing static, LLM-based, and dynamic approaches each capture only part of this problem, making them insufficient for robust real-world detection. In this paper, we present MalSkills, a neuro-symbolic framework for malicious skills detection. MalSkills first extracts security-sensitive operations from heterogeneous artifacts through a combination of symbolic parsing and LLM-assisted semantic analysis. It then constructs the skill dependency graph that links artifacts, operations, operands, and value flows across the skill. On top of this graph, MalSkills performs neuro-symbolic reasoning to infer malicious patterns or previously unseen suspicious workflows. We evaluate MalSkills on a benchmark of 200 real-world skills against 5 state-of-the-art baselines. MalSkills achieves 93% F1, outperforming the baselines by 5~87 percentage points. We further apply MalSkills to analyze 150,108 skills collected from 7 public registries, revealing 620 malicious skills. As for now, we have finished reviewing 100 of them and identified 76 previously unknown malicious skills, all of which were responsibly reported and are currently awaiting confirmation from the platforms and maintainers. These results demonstrate the potential of MalSkills in securing the agentic supply chain.


翻译:技能正日益用于扩展LLM代理,通过将提示、代码和配置打包成可复用模块。随着公共注册中心和市场的扩大,它们形成了新兴的代理供应链,但也为恶意技能引入了新的攻击面。检测恶意技能颇具挑战性,因为相关证据通常分布在异构工件中,且需结合上下文进行推理。现有的静态方法、基于LLM的方法和动态方法各自仅能捕捉该问题的部分方面,不足以实现鲁棒的实战检测。本文提出MalSkills,一个用于恶意技能检测的神经符号框架。MalSkills首先通过符号解析与LLM辅助语义分析的组合,从异构工件中提取安全敏感操作;随后构建技能依赖图,该图关联工件、操作、操作数以及跨技能的值流;在此图之上,MalSkills执行神经符号推理以推断恶意模式或先前未见过的可疑工作流。我们在包含200个真实技能的基准上对MalSkills进行评测,并对比5个最先进的基线方法。MalSkills取得了93%的F1分数,相较于基线方法提升5至87个百分点。我们进一步应用MalSkills分析从7个公共注册中心收集的150,108个技能,发现620个恶意技能。截至目前,我们已完成对其中100个技能的审查,识别出76个先前未知的恶意技能,均已负责任的报告,并正等待平台及维护者的确认。这些结果展示了MalSkills在保障代理供应链安全方面的潜力。

0
下载
关闭预览

相关内容

《基于动态图神经网络的恶意软件检测》
专知会员服务
16+阅读 · 1月28日
大型语言模型网络安全综述
专知会员服务
68+阅读 · 2024年5月12日
《使用静态污点分析检测恶意代码》CMU最新30页slides
专知会员服务
22+阅读 · 2023年10月11日
深度学习赋能的恶意代码攻防研究进展
专知会员服务
31+阅读 · 2021年4月11日
ISWC2020最佳论文《可解释假信息检测的链接可信度评价》
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
7+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
9+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
14+阅读 · 8月7日
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员