This paper presents DetCLIPv2, an efficient and scalable training framework that incorporates large-scale image-text pairs to achieve open-vocabulary object detection (OVD). Unlike previous OVD frameworks that typically rely on a pre-trained vision-language model (e.g., CLIP) or exploit image-text pairs via a pseudo labeling process, DetCLIPv2 directly learns the fine-grained word-region alignment from massive image-text pairs in an end-to-end manner. To accomplish this, we employ a maximum word-region similarity between region proposals and textual words to guide the contrastive objective. To enable the model to gain localization capability while learning broad concepts, DetCLIPv2 is trained with a hybrid supervision from detection, grounding and image-text pair data under a unified data formulation. By jointly training with an alternating scheme and adopting low-resolution input for image-text pairs, DetCLIPv2 exploits image-text pair data efficiently and effectively: DetCLIPv2 utilizes 13X more image-text pairs than DetCLIP with a similar training time and improves performance. With 13M image-text pairs for pre-training, DetCLIPv2 demonstrates superior open-vocabulary detection performance, e.g., DetCLIPv2 with Swin-T backbone achieves 40.4% zero-shot AP on the LVIS benchmark, which outperforms previous works GLIP/GLIPv2/DetCLIP by 14.4/11.4/4.5% AP, respectively, and even beats its fully-supervised counterpart by a large margin.
翻译:本文提出了DetCLIPv2,一种高效且可扩展的训练框架,通过融合大规模图像-文本对实现开放词汇目标检测(OVD)。与以往通常依赖预训练视觉语言模型(如CLIP)或通过伪标签过程利用图像-文本对的OVD框架不同,DetCLIPv2直接从海量图像-文本对中以端到端方式学习细粒度的词-区域对齐。为此,我们利用区域提议与文本词之间的最大词-区域相似度来引导对比学习目标。为使模型在获取定位能力的同时学习广泛概念,DetCLIPv2在统一数据公式下,融合了来自检测、定位和图像-文本对数据的混合监督进行训练。通过采用交替训练方案及对图像-文本对使用低分辨率输入,DetCLIPv2高效且有效地利用了图像-文本对数据:在相似训练时间内,DetCLIPv2使用的图像-文本对数量比DetCLIP多13倍,同时性能得以提升。使用1300万图像-文本对进行预训练后,DetCLIPv2展现了卓越的开放词汇检测性能,例如,采用Swin-T骨干网络的DetCLIPv2在LVIS基准上达到40.4%的零样本AP,分别比先前工作的GLIP/GLIPv2/DetCLIP高出14.4/11.4/4.5% AP,甚至以较大幅度超越了其全监督对照方法。