Motivated by the remarkable achievements of DETR-based approaches on COCO object detection and segmentation benchmarks, recent endeavors have been directed towards elevating their performance through self-supervised pre-training of Transformers while preserving a frozen backbone. Noteworthy advancements in accuracy have been documented in certain studies. Our investigation delved deeply into a representative approach, DETReg, and its performance assessment in the context of emerging models like $\mathcal{H}$-Deformable-DETR. Regrettably, DETReg proves inadequate in enhancing the performance of robust DETR-based models under full data conditions. To dissect the underlying causes, we conduct extensive experiments on COCO and PASCAL VOC probing elements such as the selection of pre-training datasets and strategies for pre-training target generation. By contrast, we employ an optimized approach named Simple Self-training which leads to marked enhancements through the combination of an improved box predictor and the Objects$365$ benchmark. The culmination of these endeavors results in a remarkable AP score of $59.3\%$ on the COCO val set, outperforming $\mathcal{H}$-Deformable-DETR + Swin-L without pre-training by $1.4\%$. Moreover, a series of synthetic pre-training datasets, generated by merging contemporary image-to-text(LLaVA) and text-to-image (SDXL) models, significantly amplifies object detection capabilities.
翻译:受基于DETR方法在COCO目标检测与分割基准上取得显著成就的启发,近期研究致力于通过自监督预训练Transformer(同时保持冻结主干的策略)来提升其性能。部分研究报告了值得关注的精度提升。我们的研究深入分析了代表性方法DETReg及其在新兴模型(如$\mathcal{H}$-Deformable-DETR)上的性能评估。遗憾的是,在全数据条件下,DETReg未能有效增强基于强健DETR模型的性能。为剖析其根本原因,我们在COCO和PASCAL VOC数据集上开展了大量实验,探究了预训练数据集选择及预训练目标生成策略等要素。相比之下,我们采用了一种名为Simple Self-training的优化方法,通过结合改进的框预测器与Objects$365$基准,实现了显著性能提升。最终,该方法在COCO验证集上取得了59.3%的惊人AP分数,超越未预训练的$\mathcal{H}$-Deformable-DETR + Swin-L达1.4%。此外,通过融合现代图像到文本(LLaVA)与文本到图像(SDXL)模型生成的一系列合成预训练数据集,进一步大幅增强了目标检测能力。