In this paper, we introduce SecQA, a novel dataset tailored for evaluating the performance of Large Language Models (LLMs) in the domain of computer security. Utilizing multiple-choice questions generated by GPT-4 based on the "Computer Systems Security: Planning for Success" textbook, SecQA aims to assess LLMs' understanding and application of security principles. We detail the structure and intent of SecQA, which includes two versions of increasing complexity, to provide a concise evaluation across various difficulty levels. Additionally, we present an extensive evaluation of prominent LLMs, including GPT-3.5-Turbo, GPT-4, Llama-2, Vicuna, Mistral, and Zephyr models, using both 0-shot and 5-shot learning settings. Our results, encapsulated in the SecQA v1 and v2 datasets, highlight the varying capabilities and limitations of these models in the computer security context. This study not only offers insights into the current state of LLMs in understanding security-related content but also establishes SecQA as a benchmark for future advancements in this critical research area.
翻译:本文提出了SecQA,一个专为评估大语言模型(LLMs)在计算机安全领域性能而设计的新颖数据集。该数据集基于《计算机系统安全:规划成功》教材,利用GPT-4生成的多项选择题,旨在评估LLMs对安全原理的理解与应用能力。我们详细阐述了SecQA的结构与设计意图,其中包含两个复杂度递增的版本,以便在不同难度层面对模型进行简洁评估。此外,我们采用0样本和5样本学习设定,对包括GPT-3.5-Turbo、GPT-4、Llama-2、Vicuna、Mistral和Zephyr在内的主流LLMs进行了全面评测。通过SecQA v1和v2数据集呈现的结果,突显了这些模型在计算机安全情境下的能力差异与局限性。本研究不仅揭示了当前LLMs在理解安全相关内容方面的状况,更将SecQA确立为该关键研究领域未来进展的基准。