Human emotion recognition plays an important role in human-computer interaction. In this paper, we present our approach to the Valence-Arousal (VA) Estimation Challenge, Expression (Expr) Classification Challenge, and Action Unit (AU) Detection Challenge of the 5th Workshop and Competition on Affective Behavior Analysis in-the-wild (ABAW). Specifically, we propose a novel multi-modal fusion model that leverages Temporal Convolutional Networks (TCN) and Transformer to enhance the performance of continuous emotion recognition. Our model aims to effectively integrate visual and audio information for improved accuracy in recognizing emotions. Our model outperforms the baseline and ranks 3 in the Expression Classification challenge.
翻译:人类情感识别在人机交互中扮演着重要角色。本文介绍了我们参与第五届野外情感行为分析研讨会与竞赛(ABAW)中效价-唤醒度(VA)估计挑战、表情(Expr)分类挑战以及动作单元(AU)检测挑战的方法。具体而言,我们提出了一种新颖的多模态融合模型,该模型利用时序卷积网络(TCN)与Transformer来提升连续情感识别的性能。我们的模型旨在有效整合视觉与音频信息,以提高情感识别的准确性。该模型优于基线方法,并在表情分类挑战中排名第三。