Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerable to evolving harms, cross-lingual failures, and cultural misalignment. Moreover, most guardian models rely on rigid, predefined safety categories that fail to generalize across diverse linguistic and sociocultural contexts. Achieving robust safety requires flexible, runtime-enforceable policies and benchmarks that reflect local norms, harm scenarios, and cultural expectations. We introduce UbuntuGuard, the first policy-based safety benchmark for African languages built from adversarial queries authored by 155 domain experts across sensitive fields, including healthcare. From these expert-crafted queries, we derive context-specific safety policies and reference responses that capture culturally grounded risk signals, enabling policy-aligned evaluation of guardian models. We evaluate 15 models, comprising seven general-purpose LLMs and eight guardian models across three distinct variants: static, dynamic, and multilingual. Our findings reveal that existing English-centric benchmarks overestimate real-world multilingual safety, cross-lingual transfer provides partial but insufficient coverage, and dynamic models, while better equipped to leverage policies at inference time, still struggle to fully localize African-language contexts. These findings highlight the urgent need for multilingual, culturally grounded safety benchmarks to enable the development of reliable and equitable guardian models for low-resource languages.
翻译:当前的守护模型主要以西方式为中心,并针对高资源语言进行了优化,导致低资源非洲语言在面对不断演变的危害、跨语言故障和文化错位时脆弱不堪。此外,大多数守护模型依赖刚性、预定义的安全类别,无法泛化到多样化的语言和社会文化背景中。实现稳健的安全需要灵活、可在运行时执行的政策和基准,以反映地方规范、危害场景和文化期望。我们提出了UbuntuGuard,这是首个基于政策的非洲语言安全基准,由155名跨敏感领域(包括医疗保健)的领域专家撰写的对抗性查询构建而成。从这些专家设计的查询中,我们推导出具有语境特定的安全政策和参考响应,以捕捉植根于文化的风险信号,从而实现对守护模型的政策对齐评估。我们评估了15个模型,包括七个通用大语言模型和八个守护模型,涵盖三种不同的变体:静态、动态和多语言。我们的发现表明,现有的以英语为中心的基准高估了现实世界中的多语言安全性,跨语言迁移仅提供部分但不充分的覆盖,而动态模型虽然在推理时更擅长利用政策,但仍难以完全本地化非洲语言语境。这些发现凸显了迫切需要多语言、植根于文化的安全基准,以推动开发针对低资源语言的可靠且公平的守护模型。