Pixel-level Video Understanding requires effectively integrating three-dimensional data in both spatial and temporal dimensions to learn accurate and stable semantic information from continuous frames. However, existing advanced models on the VSPW dataset have not fully modeled spatiotemporal relationships. In this paper, we present our solution for the PVUW competition, where we introduce masked video consistency (MVC) based on existing models. MVC enforces the consistency between predictions of masked frames where random patches are withheld. The model needs to learn the segmentation results of the masked parts through the context of images and the relationship between preceding and succeeding frames of the video. Additionally, we employed test-time augmentation, model aggeregation and a multimodal model-based post-processing method. Our approach achieves 67.27% mIoU performance on the VSPW dataset, ranking 2nd place in the PVUW2024 challenge VSS track.
翻译:像素级视频理解需要有效整合空间和时间维度上的三维数据,以从连续帧中学习准确且稳定的语义信息。然而,VSPW数据集上现有的先进模型尚未对时空关系进行充分建模。本文介绍了我们在PVUW竞赛中的解决方案,我们在现有模型基础上引入了掩码视频一致性(MVC)。MVC通过强制要求模型对随机遮蔽部分图像块后的帧进行预测时保持一致性。模型需要通过学习图像上下文以及视频前后帧之间的关系来推断被遮蔽部分的分割结果。此外,我们采用了测试时增强、模型集成以及基于多模态模型的后处理方法。我们的方法在VSPW数据集上取得了67.27%的mIoU性能,在PVUW2024挑战赛VSS赛道中排名第二。