This paper presents our approach for the VA (Valence-Arousal) estimation task in the ABAW6 competition. We devised a comprehensive model by preprocessing video frames and audio segments to extract visual and audio features. Through the utilization of Temporal Convolutional Network (TCN) modules, we effectively captured the temporal and spatial correlations between these features. Subsequently, we employed a Transformer encoder structure to learn long-range dependencies, thereby enhancing the model's performance and generalization ability. Our method leverages a multimodal data fusion approach, integrating pre-trained audio and video backbones for feature extraction, followed by TCN-based spatiotemporal encoding and Transformer-based temporal information capture. Experimental results demonstrate the effectiveness of our approach, achieving competitive performance in VA estimation on the AffWild2 dataset.
翻译:本文介绍了我们在ABAW6竞赛中针对VA(效价-唤醒度)估计任务所提出的方法。我们设计了一个综合模型,通过预处理视频帧和音频片段来提取视觉与音频特征。利用时间卷积网络(TCN)模块,我们有效捕捉了这些特征之间的时空相关性。随后,采用Transformer编码器结构学习长程依赖关系,从而提升了模型的性能与泛化能力。该方法采用多模态数据融合策略,整合预训练的音频与视频骨干网络进行特征提取,并基于TCN进行时空编码及Transformer进行时序信息捕获。实验结果表明,该方法在AffWild2数据集上取得了具有竞争力的VA估计性能。