As large language models (LLMs) are integrated into numerous applications, LLMs' safety becomes critical for both application developers and intended users. Currently, great efforts have been made to develop safety benchmarks with fine-grained taxonomies. However, these benchmarks' taxonomies are disparate with different safety policies. Thus, existing safeguards trained on these benchmarks are either coarse-grained to only distinguish between "safe'' and "unsafe,'' or constrained by the specified narrow risk taxonomies. To leverage these fine-grained safety policies across multiple safety taxonomies, we propose GSPR, a Generalizable Safety Policy Reasoner to identify unsafe inputs and outputs with violated safety taxonomies and concise explanations. Unlike prior safeguards which only cover a fixed set of risk factors, GSPR incentivizes its reasoning capability with varied safety taxonomies through reinforcement learning. Our GSPR can be trained across multiple safety benchmarks with distinct taxonomies and naturally exhibits powerful generalization ability. We conduct extensive experiments to show that GSPR significantly improves existing safety guardrails' reasoning capabilities for both safety and category prediction tasks. Moreover, GSPR also achieves the least inference token costs with explanations.
翻译:暂无翻译