Large Language Models (LLMs) are increasingly used across various domains, from software development to cyber threat intelligence. Understanding all the different fields of cybersecurity, which includes topics such as cryptography, reverse engineering, and risk assessment, poses a challenge even for human experts. To accurately test the general knowledge of LLMs in cybersecurity, the research community needs a diverse, accurate, and up-to-date dataset. To address this gap, we present CyberMetric-80, CyberMetric-500, CyberMetric-2000, and CyberMetric-10000, which are multiple-choice Q&A benchmark datasets comprising 80, 500, 2000, and 10,000 questions respectively. By utilizing GPT-3.5 and Retrieval-Augmented Generation (RAG), we collected documents, including NIST standards, research papers, publicly accessible books, RFCs, and other publications in the cybersecurity domain, to generate questions, each with four possible answers. The results underwent several rounds of error checking and refinement. Human experts invested over 200 hours validating the questions and solutions to ensure their accuracy and relevance, and to filter out any questions unrelated to cybersecurity. We have evaluated and compared 25 state-of-the-art LLM models on the CyberMetric datasets. In addition to our primary goal of evaluating LLMs, we involved 30 human participants to solve CyberMetric-80 in a closed-book scenario. The results can serve as a reference for comparing the general cybersecurity knowledge of humans and LLMs. The findings revealed that GPT-4o, GPT-4-turbo, Mixtral-8x7B-Instruct, Falcon-180B-Chat, and GEMINI-pro 1.0 were the best-performing LLMs. Additionally, the top LLMs were more accurate than humans on CyberMetric-80, although highly experienced human experts still outperformed small models such as Llama-3-8B, Phi-2 or Gemma-7b.
翻译:大语言模型(LLM)的应用日益广泛,从软件开发到网络威胁情报均有涉及。网络安全领域涵盖密码学、逆向工程和风险评估等多个主题,全面理解其所有分支领域对人类专家而言亦具挑战。为准确测试LLM在网络安全领域的通用知识水平,研究界需要一个多样化、准确且时效性强的数据集。为填补这一空白,我们提出了CyberMetric-80、CyberMetric-500、CyberMetric-2000和CyberMetric-10000,这是分别包含80、500、2000和10000道选择题的问答基准数据集。通过利用GPT-3.5和检索增强生成(RAG)技术,我们收集了网络安全领域的文档(包括NIST标准、研究论文、公开书籍、RFC及其他出版物)来生成问题,每个问题均设有四个备选答案。结果经过多轮错误检查与优化。领域专家投入超过200小时对问题与解决方案进行验证,以确保其准确性和相关性,并过滤了所有与网络安全无关的题目。我们在CyberMetric数据集上评估并比较了25个前沿LLM模型。除评估LLM的主要目标外,我们还邀请了30位人类参与者在闭卷场景下解答CyberMetric-80。该结果可为比较人类与LLM的通用网络安全知识水平提供参考。研究发现,GPT-4o、GPT-4-turbo、Mixtral-8x7B-Instruct、Falcon-180B-Chat和GEMINI-pro 1.0是表现最优的LLM。此外,在CyberMetric-80上顶尖LLM的准确率高于人类参与者,但经验丰富的人类专家仍优于Llama-3-8B、Phi-2或Gemma-7b等小型模型。