The dominant trajectory of modern machine learning has been to scale up: larger models, larger accelerators, larger memory budgets. Yet a multi-year global semiconductor supply constraint and the growing energy and carbon cost of always-online inference expose the fragility of this trajectory and motivate the opposite direction: refactoring AI and ML algorithms to fit the small, ubiquitous microcontrollers already in mass production in wearables, sensors, and edge appliances. We present an end-to-end open-source reproduction of FastGRNN, a compact gated recurrent cell, deployed on two bare-metal targets: the 8-bit Arduino (ATmega328P) and the 16-bit MSP430 (no hardware multiplier; 16 KB Flash; 512 B SRAM). Our compression pipeline combines low-rank weight factorization, iterative hard-thresholding sparsity, and per-tensor Q15 post-training quantization with explicit activation calibration. The deployed model occupies 566 bytes of weights and achieves macro F1 = 0.918 (seed 0; five-seed Q15 mean 0.853+-0.107) on the HAPT test set. It matches a PyTorch reference at 100% prediction agreement across 3,399 test windows (MCU seed 0; 99.91-100% C-equivalent across five seeds). Both platforms sustain real-time 50 Hz streaming inference (9.21 ms per sample on Arduino; 13 ms on MSP430), where a 256-entry sigmoid/tanh look-up table delivers a 30.5x speedup on the multiplier-less MSP430. Four contributions extend the original FastGRNN paper: (i) cross-platform bit-equivalent deterministic inference; (ii) characterization of recurrent warm-up latency (median 74 samples, 1.48 s; worst-case 125 samples, 2.50 s over 100 test windows); (iii) a deployable look-up-table recipe for multiplier-less embedded targets; and (iv) hardware energy characterization showing 17.7 mW active inference power, <0.09 mW idle power, and 96.7% energy reduction with the LUT.


翻译:现代机器学习的主导趋势一直是规模化扩展:更大的模型、更强的加速器、更大的内存预算。然而,全球半导体供应持续受限以及始终在线推理带来的能源和碳成本日益攀升,暴露出这一趋势的脆弱性,并推动着相反方向的发展:重构人工智能和机器学习算法,使其适配已在可穿戴设备、传感器和边缘设备中大规模量产的小型通用微控制器。我们提出一种端到端的开源复现方案,将紧凑型门控循环单元FastGRNN部署在两个裸机目标平台上:8位Arduino(ATmega328P)和16位MSP430(无硬件乘法器;16 KB闪存;512 B SRAM)。我们的压缩流水线结合了低秩权重分解、迭代硬阈值稀疏化、逐张量Q15后训练量化以及显式激活校准。部署模型仅占用566字节权重,在HAPT测试集上实现宏F1=0.918(种子0;五种子Q15均值0.853±0.107)。该模型在3,399个测试窗口上与PyTorch参考实现达到100%预测一致(MCU种子0;五种子下C等效性99.91-100%)。两个平台均支持实时50 Hz流式推理(Arduino每样本9.21 ms;MSP430每样本13 ms),其中256项sigmoid/tanh查找表在无乘法器的MSP430上实现了30.5倍加速。四项贡献扩展了原始FastGRNN论文:(i) 跨平台位等效确定性推理;(ii) 循环预热延迟特征分析(中位数74样本,1.48秒;最坏情况125样本,2.50秒,基于100个测试窗口);(iii) 面向无乘法器嵌入式目标的可部署查找表方案;(iv) 硬件能耗特征分析显示:主动推理功耗17.7 mW,空闲功耗<0.09 mW,采用查找表后能耗降低96.7%。

0
下载
关闭预览

相关内容

资源受限的大模型高效迁移学习算法研究
专知会员服务
27+阅读 · 2024年11月8日
【MIT博士论文】高效深度学习计算的模型加速
专知会员服务
34+阅读 · 2024年8月23日
【边缘智能】边缘计算驱动的深度学习加速技术
产业智能官
20+阅读 · 2019年2月8日
孟小峰:机器学习与数据库技术融合
计算机研究与发展
14+阅读 · 2018年9月6日
尽早跑通深度学习的实践代码,是入门深度学习的最快途径
算法与数据结构
22+阅读 · 2017年12月13日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Arxiv
0+阅读 · 5月13日
VIP会员
最新内容
非对称防御中的自组织临界性:俄乌战争
专知会员服务
7+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
6+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
11+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
12+阅读 · 8月8日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员