In recent years, open-vocabulary (OV) dense visual prediction (such as OV object detection, semantic, instance and panoptic segmentations) has attracted increasing research attention. However, most of existing approaches are task-specific and individually tackle each task. In this paper, we propose a Unified Open-Vocabulary Network (UOVN) to jointly address four common dense prediction tasks. Compared with separate models, a unified network is more desirable for diverse industrial applications. Moreover, OV dense prediction training data is relatively less. Separate networks can only leverage task-relevant training data, while a unified approach can integrate diverse training data to boost individual tasks. We address two major challenges in unified OV prediction. Firstly, unlike unified methods for fixed-set predictions, OV networks are usually trained with multi-modal data. Therefore, we propose a multi-modal, multi-scale and multi-task (MMM) decoding mechanism to better leverage multi-modal data. Secondly, because UOVN uses data from different tasks for training, there are significant domain and task gaps. We present a UOVN training mechanism to reduce such gaps. Experiments on four datasets demonstrate the effectiveness of our UOVN.
翻译:近年来,开放词汇(OV)密集视觉预测(如OV目标检测、语义分割、实例分割和全景分割)引起了越来越多的研究关注。然而,现有方法大多面向特定任务,各自独立解决每个任务。本文提出统一开放词汇网络(UOVN),以联合处理四种常见的密集预测任务。与独立模型相比,统一网络更适用于多样化的工业应用。此外,OV密集预测的训练数据相对较少。独立网络只能利用任务相关的训练数据,而统一方法可以整合多样化的训练数据以提升各个任务的性能。我们解决了统一OV预测中的两大挑战。首先,与固定类别预测的统一方法不同,OV网络通常使用多模态数据进行训练。因此,我们提出多模态、多尺度、多任务(MMM)解码机制,以更好地利用多模态数据。其次,由于UOVN使用来自不同任务的数据进行训练,存在显著的领域和任务差异。我们提出UOVN训练机制以减小这些差异。在四个数据集上的实验证明了UOVN的有效性。