Swear words are a common proxy to collect datasets with cyberbullying incidents. Our focus is on measuring and mitigating biases derived from spurious associations between swear words and incidents occurring as a result of such data collection strategies. After demonstrating and quantifying these biases, we introduce ID-XCB, the first data-independent debiasing technique that combines adversarial training, bias constraints and debias fine-tuning approach aimed at alleviating model attention to bias-inducing words without impacting overall model performance. We explore ID-XCB on two popular session-based cyberbullying datasets along with comprehensive ablation and generalisation studies. We show that ID-XCB learns robust cyberbullying detection capabilities while mitigating biases, outperforming state-of-the-art debiasing methods in both performance and bias mitigation. Our quantitative and qualitative analyses demonstrate its generalisability to unseen data.
翻译:脏话是收集网络欺凌事件数据集的常见代理特征。本文聚焦于测量并缓解因这种数据收集策略导致的脏话与事件之间虚假关联所产生的偏差。在证明并量化这些偏差后,我们提出了ID-XCB——首个结合对抗训练、偏差约束和去偏微调方法的数据无关去偏技术,旨在减少模型对偏置诱导词的关注而不影响整体性能。我们在两个流行的基于会话的网络欺凌数据集上探索了ID-XCB,并进行了全面的消融与泛化性研究。结果表明,ID-XCB在学习鲁棒的网络欺凌检测能力的同时缓解了偏差,在性能与偏差缓解方面均优于最先进的去偏方法。定量与定性分析证明了其对未见数据的泛化能力。