Future intelligent robots are expected to process multiple inputs simultaneously (such as image and audio data) and generate multiple outputs accordingly (such as gender and emotion), similar to humans. Recent research has shown that multi-input single-output (MISO) deep neural networks (DNN) outperform traditional single-input single-output (SISO) models, representing a significant step towards this goal. In this paper, we propose MIMONet, a novel on-device multi-input multi-output (MIMO) DNN framework that achieves high accuracy and on-device efficiency in terms of critical performance metrics such as latency, energy, and memory usage. Leveraging existing SISO model compression techniques, MIMONet develops a new deep-compression method that is specifically tailored to MIMO models. This new method explores unique yet non-trivial properties of the MIMO model, resulting in boosted accuracy and on-device efficiency. Extensive experiments on three embedded platforms commonly used in robotic systems, as well as a case study using the TurtleBot3 robot, demonstrate that MIMONet achieves higher accuracy and superior on-device efficiency compared to state-of-the-art SISO and MISO models, as well as a baseline MIMO model we constructed. Our evaluation highlights the real-world applicability of MIMONet and its potential to significantly enhance the performance of intelligent robotic systems.
翻译:未来的智能机器人有望像人类一样,能够同时处理多种输入(如图像和音频数据)并相应地生成多种输出(如性别和情绪)。近期研究表明,多输入单输出(MISO)深度神经网络(DNN)在性能上超越了传统的单输入单输出(SISO)模型,这标志着向该目标迈出了重要一步。本文提出MIMONet,一种新颖的设备端多输入多输出(MIMO)DNN框架,该框架在延迟、能耗和内存使用等关键性能指标上实现了高精度与设备端高效性。MIMONet基于现有的SISO模型压缩技术,开发了一种专为MIMO模型定制的新型深度压缩方法。该方法深入挖掘了MIMO模型独特且非平凡的特性,从而提升了模型精度与设备端效率。在机器人系统常用的三种嵌入式平台上进行的广泛实验,以及基于TurtleBot3机器人的案例研究均表明,相较于最先进的SISO与MISO模型以及我们构建的基线MIMO模型,MIMONet实现了更高的精度和更优越的设备端效率。我们的评估凸显了MIMONet在实际应用中的可行性及其显著提升智能机器人系统性能的潜力。