Automatic detection of multimodal fake news has gained a widespread attention recently. Many existing approaches seek to fuse unimodal features to produce multimodal news representations. However, the potential of powerful cross-modal contrastive learning methods for fake news detection has not been well exploited. Besides, how to aggregate features from different modalities to boost the performance of the decision-making process is still an open question. To address that, we propose COOLANT, a cross-modal contrastive learning framework for multimodal fake news detection, aiming to achieve more accurate image-text alignment. To further improve the alignment precision, we leverage an auxiliary task to soften the loss term of negative samples during the contrast process. A cross-modal fusion module is developed to learn the cross-modality correlations. An attention mechanism with an attention guidance module is implemented to help effectively and interpretably aggregate the aligned unimodal representations and the cross-modality correlations. Finally, we evaluate the COOLANT and conduct a comparative study on two widely used datasets, Twitter and Weibo. The experimental results demonstrate that our COOLANT outperforms previous approaches by a large margin and achieves new state-of-the-art results on the two datasets.
翻译:多模态虚假新闻的自动检测近年来受到广泛关注。现有方法主要通过融合单模态特征来生成多模态新闻表示,然而,强大的跨模态对比学习方法在虚假新闻检测中的潜力尚未得到充分挖掘。此外,如何聚合不同模态的特征以提升决策过程的性能仍是一个开放性问题。为此,我们提出COOLANT,一种面向多模态虚假新闻检测的跨模态对比学习框架,旨在实现更精确的图像-文本对齐。为进一步提升对齐精度,我们引入辅助任务,在对比过程中软化负样本的损失项。我们开发了跨模态融合模块以学习跨模态相关性,并实现了一种结合注意力引导模块的注意力机制,以有效且可解释地聚合对齐后的单模态表示与跨模态相关性。最后,我们在Twitter和Weibo这两个广泛使用的数据集上对COOLANT进行评估并开展对比研究。实验结果表明,我们的COOLANT方法大幅优于现有方法,在两个数据集上均取得了最新的最优结果。