Dense panoptic prediction is a key ingredient in many existing applications such as autonomous driving, automated warehouses or remote sensing. Many of these applications require fast inference over large input resolutions on affordable or even embedded hardware. We propose to achieve this goal by trading off backbone capacity for multi-scale feature extraction. In comparison with contemporaneous approaches to panoptic segmentation, the main novelties of our method are efficient scale-equivariant feature extraction, cross-scale upsampling through pyramidal fusion and boundary-aware learning of pixel-to-instance assignment. The proposed method is very well suited for remote sensing imagery due to the huge number of pixels in typical city-wide and region-wide datasets. We present panoptic experiments on Cityscapes, Vistas, COCO and the BSB-Aerial dataset. Our models outperform the state of the art on the BSB-Aerial dataset while being able to process more than a hundred 1MPx images per second on a RTX3090 GPU with FP16 precision and TensorRT optimization.
翻译:密集全景预测是众多现有应用(如自动驾驶、自动化仓储或遥感)的关键组成部分。许多此类应用要求在负担得起甚至嵌入式硬件上,对高分辨率输入实现快速推理。我们提出通过权衡骨干网络容量与多尺度特征提取来实现这一目标。与同期全景分割方法相比,我们方法的主要创新点在于:高效的尺度等变特征提取、通过金字塔融合实现的跨尺度上采样,以及像素到实例分配的边界感知学习。由于典型的城市级和区域级数据集包含海量像素,所提出的方法非常适合遥感图像。我们在Cityscapes、Vistas、COCO以及BSB-Aerial数据集上进行了全景实验。我们的模型在BSB-Aerial数据集上超越了现有技术水平,同时能在配备FP16精度和TensorRT优化的RTX3090 GPU上,每秒处理超过一百张1MPx图像。