Audio-visual deepfake detection typically employs a complementary multi-modal model to check the forgery traces in the video. These methods primarily extract forgery traces through audio-visual alignment, which results from the inconsistency between audio and video modalities. However, the traditional multi-modal forgery detection method has the problem of insufficient feature extraction and modal alignment deviation. To address this, we propose a multi-scale cross-modal transformer encoder (MSCT) for deepfake detection. Our approach includes a multi-scale self-attention to integrate the features of adjacent embeddings and a differential cross-modal attention to fuse multi-modal features. Our experiments demonstrate competitive performance on the FakeAVCeleb dataset, validating the effectiveness of the proposed structure.
翻译:音视频深度伪造检测通常采用互补的多模态模型,以检查视频中的伪造痕迹。这些方法主要通过音频与视频模态间的不一致性提取伪造痕迹。然而,传统多模态伪造检测方法存在特征提取不充分及模态对齐偏差的问题。为此,我们提出一种用于深度伪造检测的多尺度交叉模态Transformer编码器(MSCT)。该方法包含多尺度自注意力机制以整合相邻嵌入特征,以及差分交叉模态注意力机制以融合多模态特征。实验结果表明,该结构在FakeAVCeleb数据集上具有竞争力,验证了所提方法的有效性。