This paper is a summary of the work done in my PhD thesis. Where I investigate the impact of bias in NLP models on the task of hate speech detection from three perspectives: explainability, offensive stereotyping bias, and fairness. Then, I discuss the main takeaways from my thesis and how they can benefit the broader NLP community. Finally, I discuss important future research directions. The findings of my thesis suggest that the bias in NLP models impacts the task of hate speech detection from all three perspectives. And that unless we start incorporating social sciences in studying bias in NLP models, we will not effectively overcome the current limitations of measuring and mitigating bias in NLP models.
翻译:本文是对我的博士论文工作的总结。我分别从可解释性、冒犯性刻板印象偏见和公平性三个维度,探究了NLP模型中的偏见对仇恨言论检测任务的影响。随后,我讨论了论文的主要发现及其对广大NLP研究社区的潜在价值。最后,我指出了重要的未来研究方向。论文的研究结果表明,NLP模型中的偏见确实从上述三个维度影响了仇恨言论检测任务;除非我们开始引入社会科学方法来研究NLP模型中的偏见,否则将无法有效克服当前在度量和缓解NLP模型偏见方面所面临的局限。