In response to rising concerns surrounding the safety, security, and trustworthiness of Generative AI (GenAI) models, practitioners and regulators alike have pointed to AI red-teaming as a key component of their strategies for identifying and mitigating these risks. However, despite AI red-teaming's central role in policy discussions and corporate messaging, significant questions remain about what precisely it means, what role it can play in regulation, and how precisely it relates to conventional red-teaming practices as originally conceived in the field of cybersecurity. In this work, we identify recent cases of red-teaming activities in the AI industry and conduct an extensive survey of the relevant research literature to characterize the scope, structure, and criteria for AI red-teaming practices. Our analysis reveals that prior methods and practices of AI red-teaming diverge along several axes, including the purpose of the activity (which is often vague), the artifact under evaluation, the setting in which the activity is conducted (e.g., actors, resources, and methods), and the resulting decisions it informs (e.g., reporting, disclosure, and mitigation). In light of our findings, we argue that while red-teaming may be a valuable big-tent idea for characterizing a broad set of activities and attitudes aimed at improving the behavior of GenAI models, gestures towards red-teaming as a panacea for every possible risk verge on security theater. To move toward a more robust toolbox of evaluations for generative AI, we synthesize our recommendations into a question bank meant to guide and scaffold future AI red-teaming practices.
翻译:为应对生成式AI模型在安全性、可靠性和可信度方面日益增长的担忧,从业者和监管者均将AI红队测试定位为其识别并缓解这些风险策略的核心环节。然而,尽管AI红队测试在政策讨论与企业宣传中占据核心地位,其具体内涵、在监管中可发挥的作用,以及其与网络安全领域最初提出的传统红队测试实践之间的精准关系,仍存在重大疑问。本研究通过识别AI行业近期红队测试典型案例,并对相关研究文献进行广泛调研,以刻画AI红队测试实践的范畴、结构与标准。分析表明,现有AI红队测试方法与实践在多个维度存在分歧,包括活动目的(常缺乏明确性)、评估对象、实施场景(如主体、资源与方法)以及最终决策依据(如报告、披露与缓解措施)。基于研究发现,我们认为红队测试虽可作为描述旨在改善生成式AI模型行为的一系列广泛活动与态度的包容性概念,但将其奉为解决所有风险的万全之策则有流于安全示警之虞。为构建更完善的生成式AI评估工具集,我们整合研究成果形成一份问题清单,以指导并支撑未来AI红队测试实践。