As AI-generated content (AIGC) thrives, Deepfakes have expanded from single-modality falsification to cross-modal fake content creation, where either audio or visual components can be manipulated. While using two unimodal detectors can detect audio-visual deepfakes, cross-modal forgery clues could be overlooked. Existing multimodal deepfake detection methods typically establish correspondence between the audio and visual modalities for binary real/fake classification, and require the co-occurrence of both modalities. However, in real-world multi-modal applications, missing modality scenarios may occur where either modality is unavailable. In such cases, audio-visual detection methods are less practical than two independent unimodal methods. Consequently, the detector can not always obtain the number or type of manipulated modalities beforehand, necessitating a fake-modality-agnostic audio-visual detector. In this work, we propose a unified fake-modality-agnostic scenarios framework that enables the detection of multimodal deepfakes and handles missing modalities cases, no matter the manipulation hidden in audio, video, or even cross-modal forms. To enhance the modeling of cross-modal forgery clues, we choose audio-visual speech recognition (AVSR) as a preceding task, which effectively extracts speech correlation across modalities, which is difficult for deepfakes to reproduce. Additionally, we propose a dual-label detection approach that follows the structure of AVSR to support the independent detection of each modality. Extensive experiments show that our scheme not only outperforms other state-of-the-art binary detection methods across all three audio-visual datasets but also achieves satisfying performance on detection modality-agnostic audio/video fakes. Moreover, it even surpasses the joint use of two unimodal methods in the presence of missing modality cases.
翻译:随着AI生成内容(AIGC)的蓬勃发展,深度伪造已从单模态伪造扩展到跨模态虚假内容创作,其中音频或视觉成分均可能被篡改。虽然使用两个单模态检测器可以检测音视频深度伪造,但跨模态伪造线索可能被忽略。现有的多模态深度伪造检测方法通常建立音频与视觉模态之间的对应关系以进行二元真/假分类,并要求两种模态同时出现。然而,在现实多模态应用中,可能出现缺失模态场景(即某种模态不可用)。在此类情况下,音视频检测方法不如两个独立的单模态方法实用。因此,检测器无法始终预先获知被操纵模态的数量或类型,这就需要一种对伪造模态具有普适性的音视频检测器。本文提出一个统一的伪造模态无关场景框架,能够检测多模态深度伪造并处理缺失模态情况,无论篡改隐藏在音频、视频甚至跨模态形式中。为增强跨模态伪造线索的建模,我们选择音视频语音识别(AVSR)作为前置任务,该任务有效提取了跨模态的语音相关性——这是深度伪造难以复现的特征。此外,我们提出一种遵循AVSR结构的双标签检测方法,支持各模态的独立检测。大量实验表明,我们的方案不仅在三个音视频数据集上均优于其他最先进的二元检测方法,而且在检测与模态无关的音频/视频伪造方面也取得了令人满意的性能。更重要的是,在缺失模态场景下,其性能甚至超过两种单模态方法的联合使用。