Video content is rich in semantics and has the ability to evoke various emotions in viewers. In recent years, with the rapid development of affective computing and the explosive growth of visual data, affective video content analysis (AVCA) as an essential branch of affective computing has become a widely researched topic. In this study, we comprehensively review the development of AVCA over the past decade, particularly focusing on the most advanced methods adopted to address the three major challenges of video feature extraction, expression subjectivity, and multimodal feature fusion. We first introduce the widely used emotion representation models in AVCA and describe commonly used datasets. We summarize and compare representative methods in the following aspects: (1) unimodal AVCA models, including facial expression recognition and posture emotion recognition; (2) multimodal AVCA models, including feature fusion, decision fusion, and attention-based multimodal models; (3) model performance evaluation standards. Finally, we discuss future challenges and promising research directions, such as emotion recognition and public opinion analysis, human-computer interaction, and emotional intelligence.
翻译:视频内容丰富语义,能引发观者多种情感。近年来,随着情感计算的快速发展和视觉数据的爆炸式增长,情感视频内容分析(AVCA)作为情感计算的重要分支,已成为广泛研究的课题。本研究全面回顾了过去十年AVCA的发展,尤其聚焦于应对视频特征提取、表达主观性和多模态特征融合三大挑战的最先进方法。我们首先介绍AVCA中广泛使用的情感表示模型,并描述常用数据集。从以下方面总结并比较代表性方法:(1)单模态AVCA模型,包括面部表情识别和姿态情感识别;(2)多模态AVCA模型,包括特征融合、决策融合及基于注意力的多模态模型;(3)模型性能评估标准。最后,讨论未来挑战与有前景的研究方向,如情感识别与舆情分析、人机交互及情感智能。