To address the limitations of current hate speech detection models, we introduce \textsf{SGHateCheck}, a novel framework designed for the linguistic and cultural context of Singapore and Southeast Asia. It extends the functional testing approach of HateCheck and MHC, employing large language models for translation and paraphrasing into Singapore's main languages, and refining these with native annotators. \textsf{SGHateCheck} reveals critical flaws in state-of-the-art models, highlighting their inadequacy in sensitive content moderation. This work aims to foster the development of more effective hate speech detection tools for diverse linguistic environments, particularly for Singapore and Southeast Asia contexts.
翻译:为解决当前仇恨言论检测模型的局限性,我们提出\textsf{SGHateCheck}——一个专为新加坡及东南亚语言文化背景设计的新型框架。该框架扩展了HateCheck与MHC的功能测试方法,利用大语言模型将测试内容翻译及改写为新加坡主要语言,并通过本地标注员进行精炼。\textsf{SGHateCheck}揭示了现有最先进模型的重大缺陷,突显其在敏感内容审核中的不足。本研究旨在推动开发更有效的仇恨言论检测工具,以适配多元语言环境,尤其针对新加坡及东南亚地区场景。