Content moderation at scale faces the challenge of considering local cultural distinctions when assessing content. While global policies aim to maintain decision-making consistency and prevent arbitrary rule enforcement, they often overlook regional variations in interpreting natural language as expressed in content. In this study, we are looking into how moderation systems can tackle this issue by adapting to local comprehension nuances. We train large language models on extensive datasets of media news and articles to create culturally attuned models. The latter aim to capture the nuances of communication across geographies with the goal of recognizing cultural and societal variations in what is considered offensive content. We further explore the capability of these models to generate explanations for instances of content violation, aiming to shed light on how policy guidelines are perceived when cultural and societal contexts change. We find that training on extensive media datasets successfully induced cultural awareness and resulted in improvements in handling content violations on a regional basis. Additionally, these advancements include the ability to provide explanations that align with the specific local norms and nuances as evidenced by the annotators' preference in our conducted study. This multifaceted success reinforces the critical role of an adaptable content moderation approach in keeping pace with the ever-evolving nature of the content it oversees.
翻译:大规模内容审核面临在评估内容时需考虑当地文化差异的挑战。尽管全球性政策旨在保持决策一致性并防止任意规则执行,但这些政策往往忽略了对自然语言表达中内容解读的地域差异。本研究探讨审核系统如何通过适应本地理解细微差异来解决这一问题。我们利用大量媒体新闻和文章数据集训练大型语言模型,构建具备文化适应性的模型。这些模型旨在捕捉跨地域沟通的细微差别,以识别不同文化和社会背景下被认为是冒犯性内容的差异。我们进一步探索这些模型针对内容违规实例生成解释的能力,旨在揭示当文化和社会背景变化时政策指南如何被解读。研究发现,在广泛的媒体数据集上训练成功诱导了文化意识,并显著提升了区域层面处理内容违规的表现。此外,这些进展包括提供与特定本地规范和细微差别一致的解释,这在我们研究中标注者的偏好中得到了证实。这一多维度成功强化了适应性内容审核方法在跟上其监管内容不断演变性质方面的关键作用。