While speech emotion recognition (SER) research has made significant progress, achieving generalization across various corpora continues to pose a problem. We propose a novel domain adaptation technique that embodies a multitask framework with SER as the primary task, and contrastive learning and information maximisation loss as auxiliary tasks, underpinned by fine-tuning of transformers pre-trained on large language models. Empirical results obtained through experiments on well-established datasets like IEMOCAP and MSP-IMPROV, illustrate that our proposed model achieves state-of-the-art performance in SER within cross-corpus scenarios.
翻译:尽管语音情感识别(SER)研究已取得显著进展,但在不同语料库间实现泛化仍是一个难题。我们提出一种新颖的领域自适应技术,该技术以SER为主任务,以对比学习和信息最大化损失为辅助任务,构建多任务框架,并基于在大语言模型上预训练的Transformer进行微调。通过在IEMOCAP和MSP-IMPROV等公认数据集上的实验结果证明,我们提出的模型在跨语料库场景下的SER中实现了最先进的性能。