The moderation of content on online platforms is usually non-transparent. On Wikipedia, however, this discussion is carried out publicly and the editors are encouraged to use the content moderation policies as explanations for making moderation decisions. Currently, only a few comments explicitly mention those policies -- 20% of the English ones, but as few as 2% of the German and Turkish comments. To aid in this process of understanding how content is moderated, we construct a novel multilingual dataset of Wikipedia editor discussions along with their reasoning in three languages. The dataset contains the stances of the editors (keep, delete, merge, comment), along with the stated reason, and a content moderation policy, for each edit decision. We demonstrate that stance and corresponding reason (policy) can be predicted jointly with a high degree of accuracy, adding transparency to the decision-making process. We release both our joint prediction models and the multilingual content moderation dataset for further research on automated transparent content moderation.
翻译:在线平台的内容审核通常缺乏透明度。然而,在维基百科上,此类讨论以公开方式进行,编辑者被鼓励使用内容审核政策来解释其审核决策。目前,仅有少数评论明确提及这些政策——英语评论中约占20%,而德语和土耳其语评论中低至2%。为辅助理解内容审核过程,我们构建了一个新颖的多语言维基百科编辑讨论数据集,涵盖三种语言的推理依据。该数据集包含编辑者的立场(保留、删除、合并、评论),以及每个编辑决策的陈述原因和内容审核政策。我们证明,立场和相应的原因(政策)可以联合预测,且达到较高准确率,从而为决策过程增添透明度。我们公开发布了联合预测模型及多语言内容审核数据集,以促进自动化透明内容审核的进一步研究。