This paper discusses the benefits of incorporating multimodal data for improving latent emotion recognition accuracy, focusing on micro-expression (ME) and physiological signals (PS). The proposed approach presents a novel multimodal learning framework that combines ME and PS, including a 1D separable and mixable depthwise inception network, a standardised normal distribution weighted feature fusion method, and depth/physiology guided attention modules for multimodal learning. Experimental results show that the proposed approach outperforms the benchmark method, with the weighted fusion method and guided attention modules both contributing to enhanced performance.
翻译:本文探讨了融合多模态数据对于提升潜在情绪识别精度的益处,重点聚焦于微表情(ME)和生理信号(PS)。所提出的方法构建了一种新颖的多模态学习框架,该框架结合了ME与PS,包括一维可分离可混合深度初始网络、标准化正态分布加权特征融合方法,以及用于多模态学习的深度/生理引导注意力模块。实验结果表明,所提方法优于基准方法,其中加权融合方法与引导注意力模块均对性能提升有所贡献。