In this paper, we present our solutions for the 5th Workshop and Competition on Affective Behavior Analysis in-the-wild (ABAW), which includes four sub-challenges of Valence-Arousal (VA) Estimation, Expression (Expr) Classification, Action Unit (AU) Detection and Emotional Reaction Intensity (ERI) Estimation. The 5th ABAW competition focuses on facial affect recognition utilizing different modalities and datasets. In our work, we extract powerful audio and visual features using a large number of sota models. These features are fused by Transformer Encoder and TEMMA. Besides, to avoid the possible impact of large dimensional differences between various features, we design an Affine Module to align different features to the same dimension. Extensive experiments demonstrate that the superiority of the proposed method. For the VA Estimation sub-challenge, our method obtains the mean Concordance Correlation Coefficient (CCC) of 0.6066. For the Expression Classification sub-challenge, the average F1 Score is 0.4055. For the AU Detection sub-challenge, the average F1 Score is 0.5296. For the Emotional Reaction Intensity Estimation sub-challenge, the average pearson's correlations coefficient on the validation set is 0.3968. All of the results of four sub-challenges outperform the baseline with a large margin.
翻译:本文介绍了我们在第五届野外情感行为分析研讨会与竞赛(ABAW)中的解决方案,该竞赛包含四个子挑战:效价-唤醒度(VA)估计、表情(Expr)分类、动作单元(AU)检测和情绪反应强度(ERI)估计。第五届ABAW竞赛专注于利用不同模态和数据集的面部情感识别。在我们的工作中,我们使用大量最先进模型提取强大的音频和视觉特征,这些特征通过Transformer编码器和TEMMA进行融合。此外,为避免不同特征间维度差异较大可能带来的影响,我们设计了仿射模块(Affine Module)将不同特征对齐到相同维度。大量实验证明了所提方法的优越性。在效价-唤醒度估计子挑战中,我们的方法获得了0.6066的平均一致性相关系数(CCC);在表情分类子挑战中,平均F1分数为0.4055;在动作单元检测子挑战中,平均F1分数为0.5296;在情绪反应强度估计子挑战中,验证集上的平均皮尔逊相关系数为0.3968。四个子挑战的所有结果均大幅优于基线方法。