Panoramic videos contain richer spatial information and have attracted tremendous amounts of attention due to their exceptional experience in some fields such as autonomous driving and virtual reality. However, existing datasets for video segmentation only focus on conventional planar images. To address the challenge, in this paper, we present a panoramic video dataset, PanoVOS. The dataset provides 150 videos with high video resolutions and diverse motions. To quantify the domain gap between 2D planar videos and panoramic videos, we evaluate 15 off-the-shelf video object segmentation (VOS) models on PanoVOS. Through error analysis, we found that all of them fail to tackle pixel-level content discontinues of panoramic videos. Thus, we present a Panoramic Space Consistency Transformer (PSCFormer), which can effectively utilize the semantic boundary information of the previous frame for pixel-level matching with the current frame. Extensive experiments demonstrate that compared with the previous SOTA models, our PSCFormer network exhibits a great advantage in terms of segmentation results under the panoramic setting. Our dataset poses new challenges in panoramic VOS and we hope that our PanoVOS can advance the development of panoramic segmentation/tracking.
翻译:全景视频包含更丰富的空间信息,因其在自动驾驶和虚拟现实等领域提供的卓越体验而受到广泛关注。然而,现有视频分割数据集仅聚焦于常规平面图像。为解决这一难题,本文提出全景视频数据集PanoVOS。该数据集包含150个高分辨率且运动多样的视频片段。为量化二维平面视频与全景视频之间的领域差距,我们基于PanoVOS评估了15种现成的视频对象分割(VOS)模型。通过误差分析发现,所有模型均无法处理全景视频中像素级内容不连续的问题。为此,我们提出全景空间一致性Transformer(PSCFormer),该模型能有效利用前一帧的语义边界信息与当前帧进行像素级匹配。大量实验表明,与先前最先进的模型相比,我们的PSCFormer网络在全景场景下展现出了显著的分割优势。本数据集为全景视频分割提出了新挑战,期望PanoVOS能推动全景分割/追踪领域的发展。