The task of emotion recognition in conversations (ERC) benefits from the availability of multiple modalities, as provided, for example, in the video-based Multimodal EmotionLines Dataset (MELD). However, only a few research approaches use both acoustic and visual information from the MELD videos. There are two reasons for this: First, label-to-video alignments in MELD are noisy, making those videos an unreliable source of emotional speech data. Second, conversations can involve several people in the same scene, which requires the localisation of the utterance source. In this paper, we introduce MELD with Fixed Audiovisual Information via Realignment (MELD-FAIR) by using recent active speaker detection and automatic speech recognition models, we are able to realign the videos of MELD and capture the facial expressions from speakers in 96.92% of the utterances provided in MELD. Experiments with a self-supervised voice recognition model indicate that the realigned MELD-FAIR videos more closely match the transcribed utterances given in the MELD dataset. Finally, we devise a model for emotion recognition in conversations trained on the realigned MELD-FAIR videos, which outperforms state-of-the-art models for ERC based on vision alone. This indicates that localising the source of speaking activities is indeed effective for extracting facial expressions from the uttering speakers and that faces provide more informative visual cues than the visual features state-of-the-art models have been using so far. The MELD-FAIR realignment data, and the code of the realignment procedure and of the emotional recognition, are available at https://github.com/knowledgetechnologyuhh/MELD-FAIR.
翻译:对话情感识别任务得益于多模态信息的可用性,例如基于视频的多模态情感对话数据集(MELD)提供了此类信息。然而,仅有少数研究方法同时使用MELD视频中的声学与视觉信息。原因有二:首先,MELD中标签与视频的对齐存在噪声,使得这些视频作为情感语音数据的来源不可靠;其次,对话可能涉及同一场景中的多人,这需要定位话语来源。本文中,我们通过利用最新的活跃说话人检测和自动语音识别模型,引入了基于重对齐固定视听信息的MELD数据集(MELD-FAIR),成功重新对齐MELD视频,并从MELD提供的96.92%话语中捕捉到说话人的面部表情。使用自监督语音识别模型的实验表明,重对齐后的MELD-FAIR视频与MELD数据集中转录的话语更为匹配。最后,我们设计了一个基于重对齐MELD-FAIR视频训练的对话情感识别模型,其在仅依赖视觉的ERC任务上超越了现有最优模型。这表明,定位说话活动的来源确实能有效提取说话人的面部表情,且面部表情比当前最优模型所使用的视觉特征提供了更具信息量的视觉线索。MELD-FAIR重对齐数据、重对齐流程及情感识别代码均可在https://github.com/knowledgetechnologyuhh/MELD-FAIR 获取。