Facial Expression Recognition (FER) plays an important role in human-computer interactions and is used in a wide range of applications. Convolutional Neural Networks (CNN) have shown promise in their ability to classify human facial expressions, however, large CNNs are not well-suited to be implemented on resource- and energy-constrained IoT devices. In this work, we present a hierarchical framework for developing and optimizing hardware-aware CNNs tuned for deployment at the edge. We perform a comprehensive analysis across various edge AI accelerators including NVIDIA Jetson Nano, Intel Neural Compute Stick, and Coral TPU. Using the proposed strategy, we achieved a peak accuracy of 99.49% when testing on the CK+ facial expression recognition dataset. Additionally, we achieved a minimum inference latency of 0.39 milliseconds and a minimum power consumption of 0.52 Watts.
翻译:面部表情识别(Facial Expression Recognition, FER)在人机交互中发挥重要作用,并广泛应用于各类场景。卷积神经网络(CNN)在人类面部表情分类方面展现出潜力,然而大型CNN难以适用于资源与能耗受限的物联网设备。本研究提出一种分层框架,用于开发并优化面向边缘端部署的硬件感知型CNN。我们对包括NVIDIA Jetson Nano、Intel Neural Compute Stick和Coral TPU在内的多种边缘AI加速器进行了全面分析。采用所提策略,我们在CK+面部表情识别数据集上实现了99.49%的峰值准确率。此外,最低推理延迟达到0.39毫秒,最低功耗为0.52瓦特。