Facial Expression Recognition (FER) is an active research domain that has shown great progress recently, notably thanks to the use of large deep learning models. However, such approaches are particularly energy intensive, which makes their deployment difficult for edge devices. To address this issue, Spiking Neural Networks (SNNs) coupled with event cameras are a promising alternative, capable of processing sparse and asynchronous events with lower energy consumption. In this paper, we establish the first use of event cameras for FER, named "Event-based FER", and propose the first related benchmarks by converting popular video FER datasets to event streams. To deal with this new task, we propose "Spiking-FER", a deep convolutional SNN model, and compare it against a similar Artificial Neural Network (ANN). Experiments show that the proposed approach achieves comparable performance to the ANN architecture, while consuming less energy by orders of magnitude (up to 65.39x). In addition, an experimental study of various event-based data augmentation techniques is performed to provide insights into the efficient transformations specific to event-based FER.
翻译:面部表情识别(FER)是一个活跃的研究领域,近期取得了显著进展,尤其是得益于大型深度学习模型的应用。然而,这类方法特别耗能,使其难以在边缘设备上部署。为解决这一问题,结合事件相机的脉冲神经网络(SNNs)成为一种有前景的替代方案,能够处理稀疏的异步事件且能耗更低。本文首次将事件相机用于FER,命名为“基于事件的FER”,并通过将流行的视频FER数据集转换为事件流,提出了首个相关基准。为应对这一新任务,我们提出了“Spiking-FER”——一种深度卷积SNN模型,并将其与相似的人工神经网络(ANN)进行对比。实验表明,所提方法的性能与ANN架构相当,同时耗能低数个数量级(最高达65.39倍)。此外,我们还对各种基于事件的数据增强技术进行了实验研究,以提供针对基于事件的FER的高效变换的见解。