Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved. To alleviate this, we present Adaptive Language-guided Multimodal Transformer (ALMT), which incorporates an Adaptive Hyper-modality Learning (AHL) module to learn an irrelevance/conflict-suppressing representation from visual and audio features under the guidance of language features at different scales. With the obtained hyper-modality representation, the model can obtain a complementary and joint representation through multimodal fusion for effective MSA. In practice, ALMT achieves state-of-the-art performance on several popular datasets (e.g., MOSI, MOSEI and CH-SIMS) and an abundance of ablation demonstrates the validity and necessity of our irrelevance/conflict suppression mechanism.
翻译:尽管多模态情感分析(MSA)通过利用来自多个来源(如语言、视频和音频)的丰富信息被证明是有效的,但跨模态的潜在情感无关信息和冲突信息可能会阻碍性能的进一步提升。为解决这一问题,我们提出了一种自适应语言引导的多模态Transformer(ALMT),其中包含一个自适应超模态学习模块(AHL),该模块在语言特征不同尺度的引导下,从视觉和音频特征中学习抑制无关/冲突信息的表示。通过获得的超模态表示,模型能够通过多模态融合获得互补的联合表示,从而实现有效的MSA。在实际应用中,ALMT在多个流行数据集(如MOSI、MOSEI和CH-SIMS)上取得了最先进的性能,且大量消融实验验证了我们提出的无关/冲突信息抑制机制的有效性和必要性。