In this paper, we propose a sentiment-enriched lightweight network SeLiNet and an end-to-end on-device pipeline for contextual emotion recognition in images. SeLiNet model consists of body feature extractor, image aesthetics feature extractor, and learning-based fusion network which jointly estimates discrete emotion and human sentiments tasks. On the EMOTIC dataset, the proposed approach achieves an Average Precision (AP) score of 27.17 in comparison to the baseline AP score of 27.38 while reducing the model size by >85%. In addition, we report an on-device AP score of 26.42 with reduction in model size by >93% when compared to the baseline.
翻译:本文提出了一种情感增强的轻量级网络SeLiNet,以及用于图像中上下文情感识别的端到端设备端流水线。SeLiNet模型由人体特征提取器、图像美学特征提取器以及基于学习的融合网络组成,可联合完成离散情感与人类情绪感知任务。在EMOTIC数据集上,该方法的平均精度(AP)得分为27.17,相较于基线方法的27.38分,模型体积缩减超过85%。此外,我们报告了设备端AP得分为26.42,相较于基线方法,模型体积缩减超过93%。