We study the task of CVE-conditioned exploit generation, where a model drafts proof-of-concept (PoC) exploits given software vulnerability context. We adopt a data-centric approach, constructing a high-quality dataset via multi-stage preprocessing and introducing a scalable evaluation framework with LLM-as-judge and fine-grained rubrics. Under this unified setup, we benchmark 17 large language models across 8 evaluation criteria, providing systematic insights into their zero-shot capabilities. We further show that a compact 8B open-weight model, when fine-tuned on curated data, achieves over 42.5% improvement in exploit quality and rivals some proprietary models when combined with simple test-time rejection strategies. Our results highlight the importance of data quality, structured supervision, and evaluation design for reliable exploit generation, suggesting that these factors can be as critical as model scale in adapting LLMs to cybersecurity tasks.
翻译:我们研究了CVE条件驱动的利用生成任务,即模型根据软件漏洞上下文草拟概念验证(PoC)攻击代码。采用数据中心方法,通过多阶段预处理构建高质量数据集,并引入包含LLM作为评判者及细粒度评分标准的可扩展评估框架。在该统一设定下,我们依据8项评估标准对17种大型语言模型进行基准测试,系统揭示其零样本能力。进一步表明:经策展数据微调的紧凑型80亿参数开放权重模型,在利用质量上提升超42.5%,若结合简单测试时拒绝策略,可媲美部分专有模型。实验结果突显数据质量、结构化监督及评估设计对可靠利用生成的关键作用,表明在将LLM适配至网络安全任务时,这些因素可能与模型规模同等重要。