Scene analysis is essential for enabling autonomous systems, such as mobile robots, to operate in real-world environments. However, obtaining a comprehensive understanding of the scene requires solving multiple tasks, such as panoptic segmentation, instance orientation estimation, and scene classification. Solving these tasks given limited computing and battery capabilities on mobile platforms is challenging. To address this challenge, we introduce an efficient multi-task scene analysis approach, called EMSAFormer, that uses an RGB-D Transformer-based encoder to simultaneously perform the aforementioned tasks. Our approach builds upon the previously published EMSANet. However, we show that the dual CNN-based encoder of EMSANet can be replaced with a single Transformer-based encoder. To achieve this, we investigate how information from both RGB and depth data can be effectively incorporated in a single encoder. To accelerate inference on robotic hardware, we provide a custom NVIDIA TensorRT extension enabling highly optimization for our EMSAFormer approach. Through extensive experiments on the commonly used indoor datasets NYUv2, SUNRGB-D, and ScanNet, we show that our approach achieves state-of-the-art performance while still enabling inference with up to 39.1 FPS on an NVIDIA Jetson AGX Orin 32 GB.
翻译:场景分析对于移动机器人等自主系统在真实环境中的运行至关重要。然而,要全面理解场景需要解决多个任务,如全景分割、实例方向估计及场景分类。在移动平台计算能力和电池容量受限的约束下完成这些任务极具挑战。为此,我们提出一种名为EMSAFormer的高效多任务场景分析方法,该方法采用基于RGB-D Transformer的编码器同步执行上述任务。本方法基于此前发表的EMSANet框架,但研究表明,EMSANet中基于双CNN的编码器可替换为单一Transformer编码器。为实现这一目标,我们探究了如何在同一编码器中有效融合RGB与深度数据信息。为加速机器人硬件推理,我们提供了定制化的NVIDIA TensorRT扩展模块,使EMSAFormer方法实现高度优化。通过在NYUv2、SUNRGB-D和ScanNet等室内常用数据集上的大量实验,证明我们的方法在达到最优性能的同时,可在NVIDIA Jetson AGX Orin 32 GB上实现高达39.1 FPS的推理速度。