The widespread use of face retouching filters on short-video platforms has raised concerns about the authenticity of digital appearances and the impact of deceptive advertising. To address these issues, there is a pressing need to develop advanced face retouching techniques. However, the lack of large-scale and fine-grained face retouching datasets has been a major obstacle to progress in this field. In this paper, we introduce RetouchingFFHQ, a large-scale and fine-grained face retouching dataset that contains over half a million conditionally-retouched images. RetouchingFFHQ stands out from previous datasets due to its large scale, high quality, fine-grainedness, and customization. By including four typical types of face retouching operations and different retouching levels, we extend the binary face retouching detection into a fine-grained, multi-retouching type, and multi-retouching level estimation problem. Additionally, we propose a Multi-granularity Attention Module (MAM) as a plugin for CNN backbones for enhanced cross-scale representation learning. Extensive experiments using different baselines as well as our proposed method on RetouchingFFHQ show decent performance on face retouching detection. With the proposed new dataset, we believe there is great potential for future work to tackle the challenging problem of real-world fine-grained face retouching detection.
翻译:短视频平台上面部修图滤镜的广泛使用引发了人们对数字形象真实性及其在欺骗性广告中影响的担忧。为解决这些问题,迫切需要开发先进的面部修图检测技术。然而,缺乏大规模、细粒度的面部修图数据集一直是该领域进展的主要障碍。本文引入了RetouchingFFHQ,一个包含超过50万张条件修图图像的大规模细粒度面部修图数据集。RetouchingFFHQ因其大规模、高质量、细粒度以及可定制性,与先前数据集显著不同。通过涵盖四种典型面部修图操作类型及不同修图级别,我们将二值面部修图检测拓展为细粒度、多修图类型和多修图级别的估计问题。此外,我们提出了多粒度注意力模块(MAM)作为CNN骨干网络的插件,用于增强跨尺度表示学习。在RetouchingFFHQ上使用不同基线方法及我们提出的方法进行的大量实验表明,该方法在面部修图检测任务中表现优异。借助这一新数据集,我们相信未来工作将在解决现实世界中细粒度面部修图检测这一挑战性问题方面具有巨大潜力。