Visual saliency prediction for omnidirectional videos (ODVs) has shown great significance and necessity for omnidirectional videos to help ODV coding, ODV transmission, ODV rendering, etc.. However, most studies only consider visual information for ODV saliency prediction while audio is rarely considered despite its significant influence on the viewing behavior of ODV. This is mainly due to the lack of large-scale audio-visual ODV datasets and corresponding analysis. Thus, in this paper, we first establish the largest audio-visual saliency dataset for omnidirectional videos (AVS-ODV), which comprises the omnidirectional videos, audios, and corresponding captured eye-tracking data for three video sound modalities including mute, mono, and ambisonics. Then we analyze the visual attention behavior of the observers under various omnidirectional audio modalities and visual scenes based on the AVS-ODV dataset. Furthermore, we compare the performance of several state-of-the-art saliency prediction models on the AVS-ODV dataset and construct a new benchmark. Our AVS-ODV datasets and the benchmark will be released to facilitate future research.
翻译:全向视频的视觉显著性预测对于全向视频编码、传输、渲染等任务具有重要意义和必要性。然而,现有研究大多仅考虑全向视频的视觉信息进行显著性预测,尽管音频对全向视频的观看行为具有显著影响,却鲜少被纳入考量。这主要归因于缺乏大规模视听全向视频数据集及相应的分析研究。为此,本文首先构建了目前规模最大的全向视频视听显著性数据集(AVS-ODV),包含全向视频、音频及其对应的眼动追踪数据,涵盖静音、单声道和全环绕声道三种视频声音模态。随后,基于AVS-ODV数据集,我们分析了不同全向音频模态及视觉场景下观察者的视觉注意力行为。此外,我们对比了多种现有最优显著性预测模型在AVS-ODV数据集上的性能,并建立了新的基准测试。我们将公开AVS-ODV数据集及该基准,以推动相关领域的研究发展。