Automated content filtering and moderation is an important tool that allows online platforms to build striving user communities that facilitate cooperation and prevent abuse. Unfortunately, resourceful actors try to bypass automated filters in a bid to post content that violate platform policies and codes of conduct. To reach this goal, these malicious actors may obfuscate policy violating images (e.g. overlay harmful images by carefully selected benign images or visual patterns) to prevent machine learning models from reaching the correct decision. In this paper, we invite researchers to tackle this specific issue and present a new image benchmark. This benchmark, based on ImageNet, simulates the type of obfuscations created by malicious actors. It goes beyond ImageNet-$\textrm{C}$ and ImageNet-$\bar{\textrm{C}}$ by proposing general, drastic, adversarial modifications that preserve the original content intent. It aims to tackle a more common adversarial threat than the one considered by $\ell_p$-norm bounded adversaries. We evaluate 33 pretrained models on the benchmark and train models with different augmentations, architectures and training methods on subsets of the obfuscations to measure generalization. We hope this benchmark will encourage researchers to test their models and methods and try to find new approaches that are more robust to these obfuscations.
翻译:自动内容过滤与审核是帮助在线平台构建积极用户社区、促进合作并防止滥用的重要工具。然而,资源丰富的攻击者试图绕过自动过滤器,以发布违反平台政策和行为准则的内容。为实现这一目标,这些恶意行为者可能对违规图像进行混淆处理(例如,通过精心选择的良性图像或视觉模式覆盖有害图像),以阻止机器学习模型做出正确的决策。本文邀请研究人员解决这一特定问题,并提出一个新的图像基准。该基准基于ImageNet,模拟了恶意行为者创建的混淆类型。它超越了ImageNet-$\textrm{C}$和ImageNet-$\bar{\textrm{C}}$,通过提出保留原始内容意图的通用、剧烈且对抗性的修改,旨在应对比$\ell_p$范数约束对抗性威胁更常见的对抗性威胁。我们在该基准上评估了33个预训练模型,并通过在混淆子集上使用不同的数据增强、架构和训练方法来训练模型,以衡量其泛化能力。我们希望这一基准能够鼓励研究人员测试其模型和方法,并探索对这些混淆更具鲁棒性的新方法。