Speech emotion recognition aims to identify and analyze emotional states in target speech similar to humans. Perfect emotion recognition can greatly benefit a wide range of human-machine interaction tasks. Inspired by the human process of understanding emotions, we demonstrate that compared to quantized modeling, understanding speech content from a continuous perspective, akin to human-like comprehension, enables the model to capture more comprehensive emotional information. Additionally, considering that humans adjust their perception of emotional words in textual semantic based on certain cues present in speech, we design a novel search space and search for the optimal fusion strategy for the two types of information. Experimental results further validate the significance of this perception adjustment. Building on these observations, we propose a novel framework called Multiple perspectives Fusion Architecture Search (MFAS). Specifically, we utilize continuous-based knowledge to capture speech semantic and quantization-based knowledge to learn textual semantic. Then, we search for the optimal fusion strategy for them. Experimental results demonstrate that MFAS surpasses existing models in comprehensively capturing speech emotion information and can automatically adjust fusion strategy.
翻译:语音情感识别旨在像人类一样识别并分析目标语音中的情感状态。完美的情感识别能极大惠及广泛的人机交互任务。受人类理解情感过程的启发,我们证明:相较于量化建模,从连续视角理解语音内容(类似人类理解方式)能使模型捕获更全面的情感信息。此外,考虑到人类会根据语音中的某些线索调整对文本语义中情感词的感知,我们设计了一个新颖的搜索空间,并搜索两类信息的最优融合策略。实验结果进一步验证了这种感知调整的重要性。基于这些发现,我们提出一种名为多视角融合架构搜索(MFAS)的新框架。具体而言,我们利用连续型知识捕获语音语义,利用量化型知识学习文本语义,然后搜索两者的最优融合策略。实验结果表明,MFAS在综合捕获语音情感信息方面超越现有模型,并能自动调整融合策略。