Lightweight neural networks exchange fast inference for predictive strength. Conversely, large deep neural networks have low prediction error but incur prolonged inference times and high energy consumption on resource-constrained devices. This trade-off is unacceptable for latency-sensitive and performance-critical applications. Offloading inference tasks to a server is unsatisfactory due to the inevitable network congestion by high-dimensional data competing for limited bandwidth and leaving valuable client-side resources idle. This work demonstrates why existing methods cannot adequately address the need for high-performance inference in mobile edge computing. Then, we show how to overcome current limitations by introducing a novel training method to reduce bandwidth consumption in Machine-to-Machine communication and a generalizable design heuristic for resource-conscious compression models. We extensively evaluate our proposed method against a wide range of baselines for latency and compressive strength in an environment with asymmetric resource distribution between edge devices and servers. Despite our edge-oriented lightweight encoder, our method achieves considerably better compression rates.
翻译:摘要:轻量级神经网络牺牲了预测能力以换取快速推理。相反,大型深度神经网络预测误差较低,但在资源受限设备上会导致推理时间延长和能耗增加。这种权衡对于延迟敏感且性能关键的应用而言是不可接受的。将推理任务卸载至服务器并不理想,因为高维数据竞争有限带宽会导致不可避免的网络拥塞,同时宝贵的客户端资源闲置。本文阐明了现有方法为何无法充分满足移动边缘计算中对高性能推理的需求。随后,我们展示如何克服当前局限性:提出一种新颖的训练方法以减少机器间通信中的带宽消耗,以及一种适用于资源感知型压缩模型的通用设计启发式策略。我们在边缘设备与服务器资源分布不对称的环境中,针对延迟和压缩强度,将所提方法与多种基线进行了广泛评估。尽管我们采用了面向边缘的轻量级编码器,所提方法仍实现了显著更优的压缩率。