With the growing adoption of deep learning for on-device TinyML applications, there has been an ever-increasing demand for efficient neural network backbones optimized for the edge. Recently, the introduction of attention condenser networks have resulted in low-footprint, highly-efficient, self-attention neural networks that strike a strong balance between accuracy and speed. In this study, we introduce a faster attention condenser design called double-condensing attention condensers that allow for highly condensed feature embeddings. We further employ a machine-driven design exploration strategy that imposes design constraints based on best practices for greater efficiency and robustness to produce the macro-micro architecture constructs of the backbone. The resulting backbone (which we name AttendNeXt) achieves significantly higher inference throughput on an embedded ARM processor when compared to several other state-of-the-art efficient backbones (>10x faster than FB-Net C at higher accuracy and speed and >10x faster than MobileOne-S1 at smaller size) while having a small model size (>1.37x smaller than MobileNetv3-L at higher accuracy and speed) and strong accuracy (1.1% higher top-1 accuracy than MobileViT XS on ImageNet at higher speed). These promising results demonstrate that exploring different efficient architecture designs and self-attention mechanisms can lead to interesting new building blocks for TinyML applications.
翻译:随着深度学习在设备端TinyML应用中的日益普及,针对边缘端优化的高效神经网络骨干需求持续增长。近期提出的注意力凝集器网络以低开销、高效率的自注意力机制,在准确率与速度之间实现了出色平衡。本研究提出一种更快的注意力凝集器设计——双压缩注意力凝集器,其能实现高度压缩的特征嵌入。我们进一步采用机器驱动的设计探索策略,基于最佳实践施加设计约束以提升效率与鲁棒性,最终构建出该骨干网络的宏观-微观架构。所提出的骨干网络(命名为AttendNeXt)在嵌入式ARM处理器上取得了显著高于多个当前最优效率型骨干的推理吞吐量(在更高准确率和速度下比FB-Net C快10倍以上,在更小尺寸下比MobileOne-S1快10倍以上),兼具小模型体积(在更高准确率和速度下比MobileNetv3-L小1.37倍以上)与强准确率(在ImageNet上以更高速度实现比MobileViT XS高1.1%的Top-1准确率)。这些具有前景的结果表明,探索不同的高效架构设计与自注意力机制,可为TinyML应用催生有趣的新型基础构建模块。