Nowadays, multimedia forensics faces unprecedented challenges due to the rapid advancement of multimedia generation technology thereby making Image Manipulation Localization (IML) crucial in the pursuit of truth. The key to IML lies in revealing the artifacts or inconsistencies between the tampered and authentic areas, which are evident under pixel-level features. Consequently, existing studies treat IML as a low-level vision task, focusing on allocating tampered masks by crafting pixel-level features such as image RGB noises, edge signals, or high-frequency features. However, in practice, tampering commonly occurs at the object level, and different classes of objects have varying likelihoods of becoming targets of tampering. Therefore, object semantics are also vital in identifying the tampered areas in addition to pixel-level features. This necessitates IML models to carry out a semantic understanding of the entire image. In this paper, we reformulate the IML task as a high-level vision task that greatly benefits from low-level features. Based on such an interpretation, we propose a method to enhance the Masked Autoencoder (MAE) by incorporating high-resolution inputs and a perceptual loss supervision module, which is termed Perceptual MAE (PMAE). While MAE has demonstrated an impressive understanding of object semantics, PMAE can also compensate for low-level semantics with our proposed enhancements. Evidenced by extensive experiments, this paradigm effectively unites the low-level and high-level features of the IML task and outperforms state-of-the-art tampering localization methods on all five publicly available datasets.
翻译:如今,由于多媒体生成技术的飞速发展,多媒体取证面临着前所未有的挑战,这使得图像篡改定位(IML)在追求真相的过程中变得至关重要。IML的关键在于揭示篡改区域与真实区域之间的伪影或不一致性,这些在像素级特征下清晰可见。因此,现有研究将IML视为低级视觉任务,专注于通过构建像素级特征(如图像RGB噪声、边缘信号或高频特征)来分配篡改掩码。然而,在实际中,篡改通常发生在对象级别,不同类别的对象成为篡改目标的概率各不相同。因此,除了像素级特征外,对象语义在识别篡改区域中也至关重要。这要求IML模型对整个图像进行语义理解。本文将IML任务重新定义为从低级特征中获益良多的高级视觉任务。基于这一理解,我们提出了一种通过引入高分辨率输入和感知损失监督模块来增强掩码自编码器(MAE)的方法,称之为感知MAE(PMAE)。尽管MAE已展现出对对象语义的出色理解能力,但PMAE还能通过我们提出的增强方法补偿低级语义。大量实验证明,这种范式有效统一了IML任务中的低级和高级特征,并在所有五个公开数据集上优于最先进的篡改定位方法。