Today, the internet is an integral part of our daily lives, enabling people to be more connected than ever before. However, this greater connectivity and access to information increase exposure to harmful content such as cyber-bullying and cyber-hatred. Models based on machine learning and natural language offer a way to make online platforms safer by identifying hate speech in web text autonomously. However, the main difficulty is annotating a sufficiently large number of examples to train these models. This paper uses a transfer learning technique to leverage two independent datasets jointly and builds a single representation of hate speech. We build an interpretable two-dimensional visualization tool of the constructed hate speech representation -- dubbed the Map of Hate -- in which multiple datasets can be projected and comparatively analyzed. The hateful content is annotated differently across the two datasets (racist and sexist in one dataset, hateful and offensive in another). However, the common representation successfully projects the harmless class of both datasets into the same space and can be used to uncover labeling errors (false positives). We also show that the joint representation boosts prediction performances when only a limited amount of supervision is available. These methods and insights hold the potential for safer social media and reduce the need to expose human moderators and annotators to distressing online messaging.
翻译:如今,互联网已成为我们日常生活中不可或缺的一部分,使人与人之间的联系比以往任何时候都更加紧密。然而,这种更高的连通性和信息获取渠道也增加了接触网络霸凌和网络仇恨等有害内容的风险。基于机器学习与自然语言处理的模型能够自主识别网络文本中的仇恨言论,为构建更安全的在线平台提供了途径。但主要难点在于需要标注足够数量的样本来训练这些模型。本文采用迁移学习技术,联合利用两个独立数据集,构建了仇恨言论的统一表征。我们开发了一个可解释的二维可视化工具——称为"仇恨地图"——用于展示所构建的仇恨言论表征,多个数据集可在此进行投影与对比分析。这两个数据集对仇恨内容的标注方式不同(一个数据集标注种族歧视和性别歧视,另一个数据集标注仇恨和攻击性)。但这一统一表征成功地将两个数据集中的无害类别投影到同一空间,并可被用于发现标注错误(假阳性)。我们还证明,在监督信息有限的情况下,联合表征能够提升预测性能。这些方法与见解有望促进社交媒体安全,并减少人类审核员与标注者接触令人不适的网络信息的需求。