Content moderation is the process of flagging content based on pre-defined platform rules. There has been a growing need for AI moderators to safeguard users as well as protect the mental health of human moderators from traumatic content. While prior works have focused on identifying hateful/offensive language, they are not adequate for meeting the challenges of content moderation since 1) moderation decisions are based on violation of rules, which subsumes detection of offensive speech, and 2) such rules often differ across communities which entails an adaptive solution. We propose to study the challenges of content moderation by introducing a multilingual dataset of 1.8 Million Reddit comments spanning 56 subreddits in English, German, Spanish and French. We perform extensive experimental analysis to highlight the underlying challenges and suggest related research problems such as cross-lingual transfer, learning under label noise (human biases), transfer of moderation models, and predicting the violated rule. Our dataset and analysis can help better prepare for the challenges and opportunities of auto moderation.
翻译:内容审核是根据平台预设规则对内容进行标记的过程。当前,人工智能审核员的需求日益增长,既需保护用户安全,也要避免人类审核员因接触创伤性内容而影响心理健康。尽管过往研究聚焦于识别仇恨/攻击性语言,但这些方法不足以应对内容审核的挑战,原因有二:1)审核决策基于规则违反情况,这涵盖了对攻击性言论的识别;2)不同社群的规则往往存在差异,需要自适应性解决方案。我们通过引入包含180万条Reddit评论的多语言数据集(涵盖英语、德语、西班牙语和法语共56个子版块),系统研究内容审核的挑战。通过广泛实验分析,我们揭示了潜在难点,并提出相关研究问题,包括跨语言迁移、标签噪声(人类偏见)下的学习、审核模型迁移以及违规规则预测。本数据集与分析有助于提前应对自动化审核的挑战与机遇。