Deep neural network (DNN) partition is a research problem that involves splitting a DNN into multiple parts and offloading them to specific locations. Because of the recent advancement in multi-access edge computing and edge intelligence, DNN partition has been considered as a powerful tool for improving DNN inference performance when the computing resources of edge and end devices are limited and the remote transmission of data from these devices to clouds is costly. This paper provides a comprehensive survey on the recent advances and challenges in DNN partition approaches over the cloud, edge, and end devices based on a detailed literature collection. We review how DNN partition works in various application scenarios, and provide a unified mathematical model of the DNN partition problem. We developed a five-dimensional classification framework for DNN partition approaches, consisting of deployment locations, partition granularity, partition constraints, optimization objectives, and optimization algorithms. Each existing DNN partition approache can be perfectly defined in this framework by instantiating each dimension into specific values. In addition, we suggest a set of metrics for comparing and evaluating the DNN partition approaches. Based on this, we identify and discuss research challenges that have not yet been investigated or fully addressed. We hope that this work helps DNN partition researchers by highlighting significant future research directions in this domain.
翻译:深度神经网络(DNN)分割是一个研究问题,涉及将DNN划分为多个部分并卸载至特定位置。随着多接入边缘计算和边缘智能的近期发展,当边缘与端设备的计算资源受限且数据从这些设备远程传输至云端成本高昂时,DNN分割被视为提升DNN推理性能的有力工具。本文基于详细的文献收集,对面向云、边缘及端设备的DNN分割方法的最新进展与挑战进行了全面综述。我们回顾了DNN分割在不同应用场景中的运作方式,并提出了DNN分割问题的统一数学模型。针对DNN分割方法,我们构建了一个五维分类框架,涵盖部署位置、分割粒度、分割约束、优化目标及优化算法。现有每种DNN分割方法均可通过将各维度实例化为具体值在该框架中精确界定。此外,我们建议了一套用于比较和评估DNN分割方法的指标。基于此,我们识别并讨论了尚未被研究或未完全解决的研究挑战。我们期望本工作能通过凸显该领域的重要未来研究方向,为DNN分割研究者提供帮助。