Ultra-low-resolution Infrared (IR) array sensors offer a low-cost, energy-efficient, and privacy-preserving solution for people counting, with applications such as occupancy monitoring. Previous work has shown that Deep Learning (DL) can yield superior performance on this task. However, the literature was missing an extensive comparative analysis of various efficient DL architectures for IR array-based people counting, that considers not only their accuracy, but also the cost of deploying them on memory- and energy-constrained Internet of Things (IoT) edge nodes. In this work, we address this need by comparing 6 different DL architectures on a novel dataset composed of IR images collected from a commercial 8x8 array, which we made openly available. With a wide architectural exploration of each model type, we obtain a rich set of Pareto-optimal solutions, spanning cross-validated balanced accuracy scores in the 55.70-82.70% range. When deployed on a commercial Microcontroller (MCU) by STMicroelectronics, the STM32L4A6ZG, these models occupy 0.41-9.28kB of memory, and require 1.10-7.74ms per inference, while consuming 17.18-120.43 $\mu$J of energy. Our models are significantly more accurate than a previous deterministic method (up to +39.9%), while being up to 3.53x faster and more energy efficient. Further, our models' accuracy is comparable to state-of-the-art DL solutions on similar resolution sensors, despite a much lower complexity. All our models enable continuous, real-time inference on a MCU-based IoT node, with years of autonomous operation without battery recharging.
翻译:超低分辨率红外(IR)阵列传感器为人数计数提供了一种低成本、高能效且保护隐私的解决方案,可应用于 occupancy 监测等领域。已有研究表明,深度学习(DL)可在该任务中实现优越性能。然而,现有文献缺乏对不同高效DL架构在红外阵列人数计数上的全面比较分析——这种分析不仅需考虑其精确度,还需评估其部署于内存和能源受限的物联网(IoT)边缘节点时的成本。为填补这一空白,本研究对比了6种不同的DL架构,这些架构在一种新的数据集上进行验证,该数据集由商用8×8红外阵列采集的红外图像构成,并已公开提供。通过对每种模型类型进行广泛的架构探索,我们获得了一组丰富的Pareto最优解,其交叉验证平衡精确度在55.70%–82.70%之间。当部署在STMicroelectronics的商用微控制器(MCU)STM32L4A6ZG上时,这些模型占用0.41–9.28kB内存,每次推理需1.10–7.74ms,能耗为17.18–120.43 $\mu$J。相较于之前的确定性方法,我们的模型精确度显著更高(提升高达39.9%),同时速度提升达3.53倍且能效更优。此外,尽管复杂度大幅降低,我们的模型在相似分辨率传感器上的精确度与最先进的DL解决方案相当。所有模型均支持基于MCU的IoT节点进行连续、实时推理,并可在无需电池充电的情况下自主运行数年。