In our research, we pioneer a novel approach to evaluate the effectiveness of jailbreak attacks on Large Language Models (LLMs), such as GPT-4 and LLaMa2, diverging from traditional robustness-focused binary evaluations. Our study introduces two distinct evaluation frameworks: a coarse-grained evaluation and a fine-grained evaluation. Each framework, using a scoring range from 0 to 1, offers a unique perspective, enabling a more comprehensive and nuanced evaluation of attack effectiveness and empowering attackers to refine their attack prompts with greater understanding. Furthermore, we have developed a comprehensive ground truth dataset specifically tailored for jailbreak tasks. This dataset not only serves as a crucial benchmark for our current study but also establishes a foundational resource for future research, enabling consistent and comparative analyses in this evolving field. Upon meticulous comparison with traditional evaluation methods, we discovered that our evaluation aligns with the baseline's trend while offering a more profound and detailed assessment. We believe that by accurately evaluating the effectiveness of attack prompts in the Jailbreak task, our work lays a solid foundation for assessing a wider array of similar or even more complex tasks in the realm of prompt injection, potentially revolutionizing this field.
翻译:在我们的研究中,我们开创了一种评估大语言模型(LLMs,如GPT-4和LLaMa2)越狱攻击有效性的新方法,这与传统聚焦于鲁棒性的二分类评估不同。本研究引入了两种不同的评估框架:粗粒度评估和细粒度评估。每个框架均采用0到1的评分范围,提供了独特视角,从而能够更全面、细致地评估攻击有效性,并使攻击者能够更深入地优化其攻击提示。此外,我们还专门为越狱任务构建了一个全面的真实标记数据集。该数据集不仅作为当前研究的关键基准,也为未来研究奠定了基础资源,能够在这个不断发展的领域中进行一致且可比较的分析。经过与传统评估方法的细致比较,我们发现我们的评估在保持一致趋势的同时,提供了更深入、更详细的评估结果。我们相信,通过准确评估越狱任务中攻击提示的有效性,我们的工作为评估提示注入领域中更广泛甚至更复杂的类似任务奠定了坚实基础,并可能彻底变革该领域。