Researchers have long touted a vision of the future enabled by a proliferation of internet-of-things devices, including smart sensors, homes, and cities. Increasingly, embedding intelligence in such devices involves the use of deep neural networks. However, their storage and processing requirements make them prohibitive for cheap, off-the-shelf platforms. Overcoming those requirements is necessary for enabling widely-applicable smart devices. While many ways of making models smaller and more efficient have been developed, there is a lack of understanding of which ones are best suited for particular scenarios. More importantly for edge platforms, those choices cannot be analyzed in isolation from cost and user experience. In this work, we holistically explore how quantization, model scaling, and multi-modality interact with system components such as memory, sensors, and processors. We perform this hardware/software co-design from the cost, latency, and user-experience perspective, and develop a set of guidelines for optimal system design and model deployment for the most cost-constrained platforms. We demonstrate our approach using an end-to-end, on-device, biometric user authentication system using a $20 ESP-EYE board.
翻译:研究人员长期展望了由物联网设备(包括智能传感器、智能家居和智慧城市)普及所驱动的未来愿景。在设备中嵌入智能的需求日益依赖深度神经网络技术。然而,这些网络的存储和计算需求使其难以在低成本通用平台上部署。克服这些挑战是实现广泛适用的智能设备的必要条件。尽管目前已开发出多种精简模型、提升效率的方法,但学界对不同场景下最优方案的选择仍缺乏系统性认知。对于边缘平台而言,更为关键的是这些技术选择无法脱离成本与用户体验进行独立分析。本研究从全栈视角出发,系统探究量化、模型缩放与多模态技术如何与内存、传感器及处理器等系统组件交互作用。我们基于成本、延迟和用户体验的维度开展软硬件协同设计,为成本受限平台的最佳系统架构与模型部署策略制定指导方针。最终通过一个采用20美元ESP-EYE开发板的端到端设备端生物特征用户认证系统验证了本方法的有效性。