This paper conducts a robustness audit of the safety feedback of PaLM 2 through a novel toxicity rabbit hole framework introduced here. Starting with a stereotype, the framework instructs PaLM 2 to generate more toxic content than the stereotype. Every subsequent iteration it continues instructing PaLM 2 to generate more toxic content than the previous iteration until PaLM 2 safety guardrails throw a safety violation. Our experiments uncover highly disturbing antisemitic, Islamophobic, racist, homophobic, and misogynistic (to list a few) generated content that PaLM 2 safety guardrails do not evaluate as highly unsafe.
翻译:本文通过引入一种新颖的“毒性漩涡”框架,对PaLM 2的安全反馈机制进行了鲁棒性审计。该框架从刻板印象出发,指令PaLM 2生成比该刻板印象更具毒性的内容。在每次后续迭代中,持续指令PaLM 2生成比前一轮迭代更具毒性的内容,直至PaLM 2的安全护栏触发安全违规。我们的实验发现,PaLM 2的安全护栏未能将大量极端令人不安的反犹太主义、伊斯兰恐惧症、种族主义、恐同及厌女(仅举几例)内容评估为高度不安全。