Crime in the 21st century is split into a virtual and real world. However, the former has become a global menace to people's well-being and security in the latter. The challenges it presents must be faced with unified global cooperation, and we must rely more than ever on automated yet trustworthy tools to combat the ever-growing nature of online offenses. Over 10 million child sexual abuse reports are submitted to the US National Center for Missing & Exploited Children every year, and over 80% originated from online sources. Therefore, investigation centers and clearinghouses cannot manually process and correctly investigate all imagery. In light of that, reliable automated tools that can securely and efficiently deal with this data are paramount. In this sense, the scene recognition task looks for contextual cues in the environment, being able to group and classify child sexual abuse data without requiring to be trained on sensitive material. The scarcity and limitations of working with child sexual abuse images lead to self-supervised learning, a machine-learning methodology that leverages unlabeled data to produce powerful representations that can be more easily transferred to target tasks. This work shows that self-supervised deep learning models pre-trained on scene-centric data can reach 71.6% balanced accuracy on our indoor scene classification task and, on average, 2.2 percentage points better performance than a fully supervised version. We cooperate with Brazilian Federal Police experts to evaluate our indoor classification model on actual child abuse material. The results demonstrate a notable discrepancy between the features observed in widely used scene datasets and those depicted on sensitive materials.
翻译:21世纪的犯罪已分裂为虚拟世界与现实世界。然而,前者已对后者中人们的福祉与安全构成全球性威胁。其带来的挑战必须通过全球统一合作来应对,我们比以往任何时候都更需要依赖自动化且可信赖的工具,以遏制日益增长的网络犯罪势头。美国国家失踪与受虐儿童中心每年收到超过1000万份儿童性虐待举报,其中80%以上源自网络渠道。因此,调查中心与信息交换机构无法人工处理并正确调查所有图像。鉴于此,能够安全高效处理此类数据的可靠自动化工具至关重要。在此背景下,场景识别任务通过寻找环境中的上下文线索,能够在无需使用敏感材料进行训练的情况下,对儿童性虐待数据进行分组与分类。与儿童性虐待图像相关的数据稀缺性及工作限制催生了自监督学习——这种机器学习方法利用未标记数据生成强大的表征,从而更便捷地迁移至目标任务。本研究表明,基于场景中心数据预训练的自监督深度学习模型,在室内场景分类任务中可达71.6%的平衡准确率,且平均性能比全监督版本高出2.2个百分点。我们与巴西联邦警察专家合作,在真实儿童虐待材料上评估了室内分类模型。结果显示,广泛使用的场景数据集中的特征与敏感材料中描述的特征之间存在显著差异。