The capability of doing effective forensic analysis on printed and scanned (PS) images is essential in many applications. PS documents may be used to conceal the artifacts of images which is due to the synthetic nature of images since these artifacts are typically present in manipulated images and the main artifacts in the synthetic images can be removed after the PS. Due to the appeal of Generative Adversarial Networks (GANs), synthetic face images generated with GANs models are difficult to differentiate from genuine human faces and may be used to create counterfeit identities. Additionally, since GANs models do not account for physiological constraints for generating human faces and their impact on human IRISes, distinguishing genuine from synthetic IRISes in the PS scenario becomes extremely difficult. As a result of the lack of large-scale reference IRIS datasets in the PS scenario, we aim at developing a novel dataset to become a standard for Multimedia Forensics (MFs) investigation which is available at [45]. In this paper, we provide a novel dataset made up of a large number of synthetic and natural printed IRISes taken from VIPPrint Printed and Scanned face images. We extracted irises from face images and it is possible that the model due to eyelid occlusion captured the incomplete irises. To fill the missing pixels of extracted iris, we applied techniques to discover the complex link between the iris images. To highlight the problems involved with the evaluation of the dataset's IRIS images, we conducted a large number of analyses employing Siamese Neural Networks to assess the similarities between genuine and synthetic human IRISes, such as ResNet50, Xception, VGG16, and MobileNet-v2. For instance, using the Xception network, we achieved 56.76\% similarity of IRISes for synthetic images and 92.77% similarity of IRISes for real images.
翻译:对打印扫描(PS)图像进行有效取证分析的能力在许多应用中都至关重要。PS文档可能被用于掩盖图像中因合成特性而产生的伪影,由于这些伪影通常存在于篡改图像中,而合成图像中的主要伪影在PS过程后可能被消除。随着生成对抗网络(GAN)的兴起,利用GAN模型生成的合成人脸图像已难以与真实人脸区分,可能被用于伪造身份。此外,由于GAN模型在生成人脸时未考虑生理约束及其对人类虹膜的影响,因此在PS场景下区分真实虹膜与合成虹膜变得极为困难。针对PS场景中缺乏大规模参考虹膜数据集的问题,我们致力于开发一个可作为多媒体取证(MFs)研究标准的新数据集,该数据集可在文献[45]中获取。本文提出了一种由大量合成与自然打印虹膜图像组成的新数据集,这些图像源自VIPPrint打印扫描人脸数据集。我们从人脸图像中提取虹膜,但由于眼睑遮挡,模型可能捕获到不完整的虹膜。为填补提取虹膜中的缺失像素,我们应用了技术来发现虹膜图像间的复杂关联。为突出该数据集虹膜图像评估中涉及的问题,我们采用Siamese神经网络(如ResNet50、Xception、VGG16和MobileNet-v2)开展了大量分析,以评估真实虹膜与合成虹膜的相似性。例如,使用Xception网络,合成虹膜图像的相似度为56.76%,而真实虹膜图像的相似度达到92.77%。