With the ever-growing presence of social media platforms comes the increased spread of harmful content and the need for robust hate speech detection systems. Such systems easily overfit to specific targets and keywords, and evaluating them without considering distribution shifts that might occur between train and test data overestimates their benefit. We challenge hate speech models via new train-test splits of existing datasets that rely on the clustering of models' hidden representations. We present two split variants (Subset-Sum-Split and Closest-Split) that, when applied to two datasets using four pretrained models, reveal how models catastrophically fail on blind spots in the latent space. This result generalises when developing a split with one model and evaluating it on another. Our analysis suggests that there is no clear surface-level property of the data split that correlates with the decreased performance, which underscores that task difficulty is not always humanly interpretable. We recommend incorporating latent feature-based splits in model development and release two splits via the GenBench benchmark.
翻译:随着社交媒体平台的持续扩张,有害内容的传播与日俱增,亟需构建鲁棒的仇恨言论检测系统。这类系统容易过拟合特定目标和关键词,若不考虑训练集与测试集之间可能存在的分布偏移而进行评估,会高估其实际效能。我们通过基于模型隐藏表示聚类的训练-测试集新划分方法,对现有仇恨言论模型进行挑战。提出两种划分变体(Subset-Sum-Split和Closest-Split),在四个预训练模型上应用于两个数据集时,揭示了模型在潜在空间盲点上的灾难性失效。这种结果在使用一个模型生成划分、用另一模型进行评估时具有泛化性。分析表明,不存在与性能下降相关的数据划分表层特征,这印证了任务难度并非总是符合人类直觉。我们建议在模型开发中纳入基于潜在特征的划分方法,并通过GenBench基准发布两种划分方案。