We present, to our knowledge, the most comprehensive cross-model evaluation of LLM agents on offensive cybersecurity tasks, benchmarking 10 frontier models from 7 providers on all 200 challenges of the NYU CTF Bench. Building on the D-CIPHER multi-agent framework, we extend it with multi-provider backend support, a custom Kali Linux environment with over 100 pre-installed penetration testing tools, and runtime tool-discovery agents. Through a controlled factorial study, we find that the Kali Linux environment yields a +9.5 percentage-point improvement over Ubuntu, while auto-prompting and category-specific tips often degrade performance in well-equipped environments. Among models, Claude 4.5 Opus achieves the highest solve rate (59%), followed by Gemini 3 Pro (52%), with Gemini 3 Flash offering the best cost-efficiency at $0.05 per solve. Asymmetric planner/executor model assignments provide no meaningful benefit while coherent same-model configurations consistently outperform mixed-tier pairings. Our results indicate that environment tooling and model selection emerge as the strongest drivers of performance, whereas prompt engineering interventions show diminishing or negative returns in well-equipped environments. Reported performance reflects both model reasoning ability and compatibility with agent tooling and API integration.
翻译:我们提出(据我们所知)针对大语言模型智能体在网络安全攻击任务中最全面的跨模型评估,对来自7家供应商的10个前沿模型在NYU CTF Bench的全部200个挑战中进行了基准测试。基于D-CIPHER多智能体框架,我们通过多供应商后端支持、配备超过100个预装渗透测试工具的自定义Kali Linux环境以及运行时工具发现智能体对其进行了扩展。通过控制因子研究,我们发现Kali Linux环境较Ubuntu带来+9.5个百分点的性能提升,而自动提示生成和分类提示在设备完善的环境中往往降低性能。在模型中,Claude 4.5 Opus实现最高解决率(59%),其次为Gemini 3 Pro(52%),Gemini 3 Flash以每个解决方案0.05美元的成本效率最佳。非对称规划器/执行器模型分配未带来实质收益,而连贯的同模型配置始终优于混合层级配对。结果表明环境工具和模型选择是性能的最强驱动力,而提示工程干预在设备完善的环境中呈现递减或负回报。报告性能同时反映模型推理能力及其与智能体工具和API集成的兼容性。