This paper introduces a novel approach named CrossVideo, which aims to enhance self-supervised cross-modal contrastive learning in the field of point cloud video understanding. Traditional supervised learning methods encounter limitations due to data scarcity and challenges in label acquisition. To address these issues, we propose a self-supervised learning method that leverages the cross-modal relationship between point cloud videos and image videos to acquire meaningful feature representations. Intra-modal and cross-modal contrastive learning techniques are employed to facilitate effective comprehension of point cloud video. We also propose a multi-level contrastive approach for both modalities. Through extensive experiments, we demonstrate that our method significantly surpasses previous state-of-the-art approaches, and we conduct comprehensive ablation studies to validate the effectiveness of our proposed designs.
翻译:本文提出一种名为CrossVideo的新方法,旨在增强点云视频理解领域的自监督跨模态对比学习。传统监督学习方法因数据稀缺和标签获取困难而面临局限性。为解决这些问题,我们提出一种自监督学习方法,利用点云视频与图像视频之间的跨模态关系获取有意义的特征表示。通过采用模态内和跨模态对比学习技术,促进对点云视频的有效理解。此外,我们还针对两种模态提出了一种多层级对比方法。大量实验表明,我们的方法显著超越了现有最优技术,并通过全面的消融研究验证了所提设计的有效性。