The growing interest in omnidirectional videos (ODVs) that capture the full field-of-view (FOV) has gained 360-degree saliency prediction importance in computer vision. However, predicting where humans look in 360-degree scenes presents unique challenges, including spherical distortion, high resolution, and limited labelled data. We propose a novel vision-transformer-based model for omnidirectional videos named SalViT360 that leverages tangent image representations. We introduce a spherical geometry-aware spatiotemporal self-attention mechanism that is capable of effective omnidirectional video understanding. Furthermore, we present a consistency-based unsupervised regularization term for projection-based 360-degree dense-prediction models to reduce artefacts in the predictions that occur after inverse projection. Our approach is the first to employ tangent images for omnidirectional saliency prediction, and our experimental results on three ODV saliency datasets demonstrate its effectiveness compared to the state-of-the-art.
翻译:随着捕捉全视场角(FOV)的全向视频(ODVs)受到的关注日益增长,360度显著性预测在计算机视觉领域变得愈发重要。然而,预测人类在360度场景中的注视点面临独特挑战,包括球形畸变、高分辨率以及有限的标注数据。我们提出了一种基于视觉Transformer的新型全向视频模型SalViT360,该模型利用切面图像表示。我们引入了一种球面几何感知的时空自注意力机制,能够有效理解全向视频。此外,我们提出了一种基于一致性的无监督正则化项,用于基于投影的360度密集预测模型,以减少逆投影后预测结果中出现的伪影。我们的方法是首个将切面图像应用于全向显著性预测的工作,在三个ODV显著性数据集上的实验结果表明,与现有最优方法相比,该方法具有显著优势。