Facial expression recognition is vital for human behavior analysis, and deep learning has enabled models that can outperform humans. However, it is unclear how closely they mimic human processing. This study aims to explore the similarity between deep neural networks and human perception by comparing twelve different networks, including both general object classifiers and FER-specific models. We employ an innovative global explainable AI method to generate heatmaps, revealing crucial facial regions for the twelve networks trained on six facial expressions. We assess these results both quantitatively and qualitatively, comparing them to ground truth masks based on Friesen and Ekman's description and among them. We use Intersection over Union (IoU) and normalized correlation coefficients for comparisons. We generate 72 heatmaps to highlight critical regions for each expression and architecture. Qualitatively, models with pre-trained weights show more similarity in heatmaps compared to those without pre-training. Specifically, eye and nose areas influence certain facial expressions, while the mouth is consistently important across all models and expressions. Quantitatively, we find low average IoU values (avg. 0.2702) across all expressions and architectures. The best-performing architecture averages 0.3269, while the worst-performing one averages 0.2066. Dendrograms, built with the normalized correlation coefficient, reveal two main clusters for most expressions: models with pre-training and models without pre-training. Findings suggest limited alignment between human and AI facial expression recognition, with network architectures influencing the similarity, as similar architectures prioritize similar facial regions.
翻译:面部表情识别对于人类行为分析至关重要,深度学习已使模型性能超越人类。然而,这些模型在多大程度上模拟了人类处理过程仍不清楚。本研究旨在通过比较十二种不同网络(包括通用目标分类器与特定于面部表情识别的模型)来探究深度神经网络与人类感知的相似性。我们采用一种创新的全局可解释人工智能方法生成热力图,揭示在六种面部表情任务上训练的十二个网络所关注的关键面部区域。通过定量与定性分析,我们将其与基于弗里森和埃克曼描述的真实掩码进行比较,并评估网络间的差异。采用交并比与归一化相关系数进行对比,共生成72张热力图以突出每种表情和架构的关键区域。定性分析表明,预训练模型的熱力图相似度高于无预训练模型。具体而言,眼睛和鼻子区域对某些面部表情具有显著影响,而嘴巴区域在所有模型和表情中始终重要。定量分析显示,所有表情和架构的平均交并比值为0.2702,最佳与最差架构的均值分别为0.3269与0.2066。基于归一化相关系数构建的树状图揭示大多数表情存在两大聚类:预训练模型与无预训练模型。研究表明,人工智能与人类的面部表情识别存在有限对齐,网络架构影响相似性,结构相似的网络会优先关注相似的面部区域。