Existing approaches to unsupervised video instance segmentation typically rely on motion estimates and experience difficulties tracking small or divergent motions. We present VideoCutLER, a simple method for unsupervised multi-instance video segmentation without using motion-based learning signals like optical flow or training on natural videos. Our key insight is that using high-quality pseudo masks and a simple video synthesis method for model training is surprisingly sufficient to enable the resulting video model to effectively segment and track multiple instances across video frames. We show the first competitive unsupervised learning results on the challenging YouTubeVIS-2019 benchmark, achieving 50.7% APvideo^50 , surpassing the previous state-of-the-art by a large margin. VideoCutLER can also serve as a strong pretrained model for supervised video instance segmentation tasks, exceeding DINO by 15.9% on YouTubeVIS-2019 in terms of APvideo.
翻译:现有无监督视频实例分割方法通常依赖运动估计,但在跟踪微小或非刚性运动时面临困难。我们提出VideoCutLER——一种无需光流等运动信号或自然视频训练的无监督多实例视频分割简易方法。核心发现是:仅需使用高质量伪掩码和简单的视频合成方法进行模型训练,就能使视频模型有效分割并跟踪视频帧中的多个实例,效果之好令人惊讶。我们在具有挑战性的YouTubeVIS-2019基准上首次展现了具有竞争力的无监督学习结果,达到50.7% APvideo^50,大幅超越先前最先进方法。VideoCutLER还可作为监督视频实例分割任务的强预训练模型,在YouTubeVIS-2019上APvideo指标相较DINO提升15.9%。