Command-line interface (CLI) agents powered by large language models (LLMs) can interpret natural-language requests, plan multi-step tasks, execute shell commands, and modify files and system state. As these agents are increasingly used for operating-system (OS) workflows, it is important to evaluate whether they can be misused to carry out security-relevant operations. Existing benchmarks often lack an attacker-knowledge model grounded in tactics, techniques, and procedures (TTPs), provide limited coverage of end-to-end kill chains, rely on simplified single-host environments, or use LLM-as-a-judge for success evaluation. We introduce AdvCLI, an MITRE ATT&CK-aligned benchmark for evaluating OS-level misuse risks of CLI agents in a controlled multi-host sandbox. AdvCLI contains 140 tasks: 40 direct malicious requests, 74 TTP-based tasks, and 26 end-to-end kill chains. Each task is paired with deterministic hard-coded verification protocols that check whether the requested OS-level effect is realized. We evaluate seven CLI agents and products built on nine foundation models, including ReAct, OpenClaw, OpenAI Agent SDK, Claude Code, Gemini CLI, Cursor CLI, and Cursor IDE. Results show that current CLI agents frequently proceed beyond refusal and can complete a non-negligible fraction of malicious OS-level tasks, especially when requests include TTP-style attacker knowledge. AdvCLI provides a reproducible testbed for evaluating these risks and for developing stronger safety mechanisms for tool-using CLI agents in the future.


翻译:暂无翻译

0
下载
关闭预览

相关内容

综述 | 终端智能体:命令行环境中的 AI Agents
专知会员服务
9+阅读 · 8月24日
综述 | 从问答到任务完成:Agent系统与Harness设计
专知会员服务
19+阅读 · 6月24日
Agent Harness综述:大模型智能体执行器工程全景
专知会员服务
29+阅读 · 5月28日
DeepSeek 版Claude Code,免费小白安装教程来了!
专知会员服务
24+阅读 · 5月5日
OpenNRE 2.0:可一键运行的开源关系抽取工具包
PaperWeekly
22+阅读 · 2019年10月30日
GraphSAGE:我寻思GCN也没我牛逼
极市平台
12+阅读 · 2019年8月12日
微软机器阅读理解在一场多轮对话挑战中媲美人类
微软丹棱街5号
19+阅读 · 2019年5月14日
GCNet:当Non-local遇见SENet
极市平台
11+阅读 · 2019年5月9日
A Technical Overview of AI & ML in 2018 & Trends for 2019
待字闺中
18+阅读 · 2018年12月24日
专访 | Recurrent AI:呼叫系统的「变废为宝」
机器之心
12+阅读 · 2018年11月28日
Network Embedding 指南
专知
22+阅读 · 2018年8月13日
论文浅尝 | Question Answering over Freebase
开放知识图谱
19+阅读 · 2018年1月9日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
VIP会员
最新内容
《最强大的军事网状网络》
专知会员服务
2+阅读 · 9月7日
《预测陆军征兵任务分配》110页
专知会员服务
2+阅读 · 9月7日
分层反无人机系统发展新趋势
专知会员服务
10+阅读 · 9月3日
何为协作武器?
专知会员服务
10+阅读 · 9月1日
相关资讯
OpenNRE 2.0:可一键运行的开源关系抽取工具包
PaperWeekly
22+阅读 · 2019年10月30日
GraphSAGE:我寻思GCN也没我牛逼
极市平台
12+阅读 · 2019年8月12日
微软机器阅读理解在一场多轮对话挑战中媲美人类
微软丹棱街5号
19+阅读 · 2019年5月14日
GCNet:当Non-local遇见SENet
极市平台
11+阅读 · 2019年5月9日
A Technical Overview of AI & ML in 2018 & Trends for 2019
待字闺中
18+阅读 · 2018年12月24日
专访 | Recurrent AI:呼叫系统的「变废为宝」
机器之心
12+阅读 · 2018年11月28日
Network Embedding 指南
专知
22+阅读 · 2018年8月13日
论文浅尝 | Question Answering over Freebase
开放知识图谱
19+阅读 · 2018年1月9日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员