Large language models are increasingly deployed across professional domains, bringing hard-to-predict risks, including the generation of harmful or disrespectful content. Although substantial progress has been made in developing safety evaluation datasets, existing resources remain overwhelmingly English- and Chinese-centric. This limitation is particularly pronounced when evaluating languages that operate within shared sociocultural, legal, and ethical contexts. To address this gap, we introduce Schützen: a German--Bulgarian safety dataset designed to assess model answerability under risk, covering both a low-resource language (Bulgarian) and a high-resource language (German). Experiments with multilingual and language-specific LLMs reveal pronounced cross-language differences in safety behavior, highlighting the necessity of tailored, region-specific evaluation resources to support the responsible deployment of LLMs in Germany and Bulgaria. Datasets and code are available at https://github.com/xnlp-lab/Schutzen. Warning: this paper contains examples that may be offensive, harmful, or biased.
翻译:大语言模型日益广泛地部署于各类专业领域,由此带来了难以预测的风险,包括生成有害或不尊重他人的内容。尽管安全性评估数据集的开发取得了显著进展,但现有资源仍过度集中于英语和汉语语境。这一局限在评估共享社会文化、法律及伦理背景下的语言时尤为突出。为弥补这一空白,我们提出Schützen:一个面向德语-保加利亚语的安全性数据集,旨在评估模型在风险情境下的应答可靠性,覆盖低资源语言(保加利亚语)与高资源语言(德语)两类场景。基于多语言及特定语言大语言模型的实验揭示了安全行为中显著的跨语言差异,凸显了为德国和保加利亚负责任地部署大语言模型而定制区域特异性评估资源的必要性。数据集与代码参见https://github.com/xnlp-lab/Schutzen。警告:本文包含可能具有冒犯性、有害性或偏见性的示例。