Generative techniques continue to evolve at an impressively high rate, driven by the hype about these technologies. This rapid advancement severely limits the application of deepfake detectors, which, despite numerous efforts by the scientific community, struggle to achieve sufficiently robust performance against the ever-changing content. To address these limitations, in this paper, we propose an analysis of two continuous learning techniques on a Short and a Long sequence of fake media. Both sequences include a complex and heterogeneous range of deepfakes generated from GANs, computer graphics techniques, and unknown sources. Our study shows that continual learning could be important in mitigating the need for generalizability. In fact, we show that, although with some limitations, continual learning methods help to maintain good performance across the entire training sequence. For these techniques to work in a sufficiently robust way, however, it is necessary that the tasks in the sequence share similarities. In fact, according to our experiments, the order and similarity of the tasks can affect the performance of the models over time. To address this problem, we show that it is possible to group tasks based on their similarity. This small measure allows for a significant improvement even in longer sequences. This result suggests that continual techniques can be combined with the most promising detection methods, allowing them to catch up with the latest generative techniques. In addition to this, we propose an overview of how this learning approach can be integrated into a deepfake detection pipeline for continuous integration and continuous deployment (CI/CD). This allows you to keep track of different funds, such as social networks, new generative tools, or third-party datasets, and through the integration of continuous learning, allows constant maintenance of the detectors.
翻译:生成技术在这些技术热潮的推动下,正以惊人的速度持续演进。这种快速进步严重限制了深度伪造检测器的应用,尽管科学界付出了诸多努力,这些检测器仍难以在不断变化的内容面前实现足够鲁棒的性能。为应对这些局限性,本文提出在短序列和长序列伪造媒体上分析两种持续学习技术。这两个序列均包含由GAN、计算机图形学技术及未知来源生成的复杂且异质的深度伪造内容。我们的研究表明,持续学习对于减轻泛化需求可能具有重要意义。事实上,我们发现,尽管存在某些限制,持续学习方法有助于在整个训练序列中保持良好的性能。然而,要使这些技术以足够鲁棒的方式工作,序列中的任务必须具有相似性。根据我们的实验,任务的顺序和相似性确实会影响模型随时间推移的性能。针对该问题,我们证明可以基于任务相似性进行分组。这种简单措施即使在较长序列中也能带来显著改进。该结果表明,持续学习技术可与最有前景的检测方法相结合,使其能够跟上最新生成技术的发展。此外,我们概述了如何将这种学习方法集成到深度伪造检测流程中,以实现持续集成与持续部署(CI/CD)。这使得系统能够追踪不同来源(如社交网络、新型生成工具或第三方数据集),并通过集成持续学习技术,实现对检测器的持续维护。