With the rise of social media, a rise of hateful content can be observed. Even though the understanding and definitions of hate speech varies, platforms, communities, and legislature all acknowledge the problem. Therefore, adolescents are a new and active group of social media users. The majority of adolescents experience or witness online hate speech. Research in the field of automated hate speech classification has been on the rise and focuses on aspects such as bias, generalizability, and performance. To increase generalizability and performance, it is important to understand biases within the data. This research addresses the bias of youth language within hate speech classification and contributes by providing a modern and anonymized hate speech youth language data set consisting of 88.395 annotated chat messages. The data set consists of publicly available online messages from the chat platform Discord. ~6,42% of the messages were classified by a self-developed annotation schema as hate speech. For 35.553 messages, the user profiles provided age annotations setting the average author age to under 20 years old.
翻译:随着社交媒体的兴起,仇恨内容也呈现增长趋势。尽管对仇恨言论的理解和定义存在差异,但平台、社区及立法机构均承认这一问题的存在。青少年作为社交媒体用户中新兴且活跃的群体,多数都曾经历或目睹过在线仇恨言论。近年来,自动仇恨言论分类领域的研究日益增多,重点关注偏差、泛化能力和性能等维度。为提升泛化能力和性能,理解数据中的偏差至关重要。本研究聚焦仇恨言论分类中青少年语言的偏差问题,通过提供包含88,395条带注释聊天消息的现代匿名化青少年语言仇恨言论数据集作出贡献。该数据集由聊天平台Discord的公开在线消息组成,其中约6.42%的消息被自研标注方案判定为仇恨言论。在35,553条消息中,用户资料提供了年龄标注,显示消息作者平均年龄不足20岁。