With the emergence of VR and AR, 360{\deg} data attracts increasing attention from the computer vision and multimedia communities. Typically, 360{\deg} data is projected into 2D ERP (equirectangular projection) images for feature extraction. However, existing methods cannot handle the distortions that result from the projection, hindering the development of 360-data-based tasks. Therefore, in this paper, we propose a Transformer-based model called DATFormer to address the distortion problem. We tackle this issue from two perspectives. Firstly, we introduce two distortion-adaptive modules. The first is a Distortion Mapping Module, which guides the model to pre-adapt to distorted features globally. The second module is a Distortion-Adaptive Attention Block that reduces local distortions on multi-scale features. Secondly, to exploit the unique characteristics of 360{\deg} data, we present a learnable relation matrix and use it as part of the positional embedding to further improve performance. Extensive experiments are conducted on three public datasets, and the results show that our model outperforms existing 2D SOD (salient object detection) and 360 SOD methods.
翻译:随着VR和AR技术的兴起,360°数据日益受到计算机视觉和多媒体领域的关注。通常,360°数据被投影为二维ERP(等距柱状投影)图像用于特征提取。然而,现有方法无法有效处理投影导致的畸变问题,阻碍了基于360°数据任务的进一步发展。为此,本文提出一种名为DATFormer的Transformer模型以解决畸变问题。我们从两个角度入手:首先,引入两个畸变自适应模块——畸变映射模块(引导模型全局预适应畸变特征)与畸变自适应注意力模块(降低多尺度特征的局部畸变)。其次,为充分利用360°数据的独特特性,我们提出可学习的关系矩阵,并将其作为位置编码的一部分以进一步提升性能。在三个公开数据集上的大量实验表明,我们的模型优于现有二维SOD(显著目标检测)与360° SOD方法。