Recently, emotion recognition based on physiological signals has emerged as a field with intensive research. The utilization of multi-modal, multi-channel physiological signals has significantly improved the performance of emotion recognition systems, due to their complementarity. However, effectively integrating emotion-related semantic information from different modalities and capturing inter-modal dependencies remains a challenging issue. Many existing multimodal fusion methods ignore either token-to-token or channel-to-channel correlations of multichannel signals from different modalities, which limits the classification capability of the models to some extent. In this paper, we propose a comprehensive perspective of multimodal fusion that integrates channel-level and token-level cross-modal interactions. Specifically, we introduce a unified cross attention module called Token-chAnnel COmpound (TACO) Cross Attention to perform multimodal fusion, which simultaneously models channel-level and token-level dependencies between modalities. Additionally, we propose a 2D position encoding method to preserve information about the spatial distribution of EEG signal channels, then we use two transformer encoders ahead of the fusion module to capture long-term temporal dependencies from the EEG signal and the peripheral physiological signal, respectively. Subject-independent experiments on emotional dataset DEAP and Dreamer demonstrate that the proposed model achieves state-of-the-art performance.
翻译:近年来,基于生理信号的情感识别已成为一个研究热点。由于多模态、多通道生理信号具有互补性,其应用显著提升了情感识别系统的性能。然而,如何有效整合来自不同模态的情感语义信息并捕捉模态间依赖关系仍是一个具有挑战性的问题。现有许多多模态融合方法忽视了不同模态多通道信号中令牌-令牌或通道-通道的关联性,这在一定程度上限制了模型的分类能力。本文提出了一种融合通道级与令牌级跨模态交互的多模态融合综合视角。具体而言,我们引入了一种名为令牌-通道复合交叉注意力(Token-chAnnel COmpound, TACO)的统一交叉注意力模块用于多模态融合,该模块可同时建模模态间的通道级与令牌级依赖关系。此外,我们提出了一种二维位置编码方法以保留脑电信号通道的空间分布信息,并在融合模块前分别采用两个Transformer编码器捕获脑电信号与外围生理信号的长期时序依赖。基于情感数据集DEAP和Dreamer的受试者独立实验表明,所提模型达到了当前最优性能。