Fusing multiple modalities has proven effective for multimodal information processing. However, the incongruity between modalities poses a challenge for multimodal fusion, especially in affect recognition. In this study, we first analyze how the salient affective information in one modality can be affected by the other, and demonstrate that inter-modal incongruity exists latently in crossmodal attention. Based on this finding, we propose the Hierarchical Crossmodal Transformer with Dynamic Modality Gating (HCT-DMG), a lightweight incongruity-aware model, which dynamically chooses the primary modality in each training batch and reduces fusion times by leveraging the learned hierarchy in the latent space to alleviate incongruity. The experimental evaluation on five benchmark datasets: CMU-MOSI, CMU-MOSEI, and IEMOCAP (sentiment and emotion), where incongruity implicitly lies in hard samples, as well as UR-FUNNY (humour) and MUStaRD (sarcasm), where incongruity is common, verifies the efficacy of our approach, showing that HCT-DMG: 1) outperforms previous multimodal models with a reduced size of approximately 0.8M parameters; 2) recognizes hard samples where incongruity makes affect recognition difficult; 3) mitigates the incongruity at the latent level in crossmodal attention.
翻译:融合多模态已被证明对多模态信息处理有效。然而,模态间的不协调对多模态融合构成挑战,尤其在情感识别中。本研究首先分析一种模态中的显著情感信息如何受到另一种模态的影响,并证明跨模态注意力中潜藏着模态间的不协调性。基于此发现,我们提出带有动态模态门控的层级跨模态Transformer(HCT-DMG)——一种轻量级不协调感知模型。该模型在每个训练批次中动态选择主模态,并通过利用隐空间中的学习层级减少融合次数以缓解不协调。在五个基准数据集上的实验评估:CMU-MOSI、CMU-MOSEI和IEMOCAP(情绪与情感,其中不协调隐式存在于困难样本中),以及UR-FUNNY(幽默)和MUStaRD(讽刺,其中不协调普遍存在),验证了本方法的有效性。结果表明HCT-DMG:1)以约0.8M参数的缩小规模超越以往多模态模型;2)能识别因不协调导致情感识别困难的困难样本;3)在跨模态注意力的隐层水平上缓解不协调性。