This paper conducts a robustness audit of the safety feedback of PaLM 2 through a novel toxicity rabbit hole framework introduced here. Starting with a stereotype, the framework instructs PaLM 2 to generate more toxic content than the stereotype. Every subsequent iteration it continues instructing PaLM 2 to generate more toxic content than the previous iteration until PaLM 2 safety guardrails throw a safety violation. Our experiments uncover highly disturbing antisemitic, Islamophobic, racist, homophobic, and misogynistic (to list a few) generated content that PaLM 2 safety guardrails do not evaluate as highly unsafe.
翻译:本文通过引入一种新颖的毒性兔子洞框架,对PaLM 2的安全反馈进行了鲁棒性审计。该框架从一个刻板印象开始,指示PaLM 2生成比该刻板印象更具毒性的内容。在随后的每一次迭代中,它持续指示PaLM 2生成比前一次迭代更具毒性的内容,直至PaLM 2的安全护栏判定为安全违规。我们的实验揭示了PaLM 2安全护栏未评估为高度不安全的高度令人不安的反犹太、仇视伊斯兰、种族主义、恐同以及厌女(仅举几例)生成内容。