The volume and diversity of training data are critical for modern deep learningbased methods. Compared to the massive amount of labeled perspective images, 360 panoramic images fall short in both volume and diversity. In this paper, we propose PanoMixSwap, a novel data augmentation technique specifically designed for indoor panoramic images. PanoMixSwap explicitly mixes various background styles, foreground furniture, and room layouts from the existing indoor panorama datasets and generates a diverse set of new panoramic images to enrich the datasets. We first decompose each panoramic image into its constituent parts: background style, foreground furniture, and room layout. Then, we generate an augmented image by mixing these three parts from three different images, such as the foreground furniture from one image, the background style from another image, and the room structure from the third image. Our method yields high diversity since there is a cubical increase in image combinations. We also evaluate the effectiveness of PanoMixSwap on two indoor scene understanding tasks: semantic segmentation and layout estimation. Our experiments demonstrate that state-of-the-art methods trained with PanoMixSwap outperform their original setting on both tasks consistently.
翻译:训练数据的数量和多样性对于基于现代深度学习的方法至关重要。与大量标记的透视图像相比,360度全景图像在数量和多样性方面均显不足。本文提出PanoMixSwap,一种专为室内全景图像设计的新型数据增强技术。PanoMixSwap通过显式混合现有室内全景数据集中的多种背景风格、前景家具和房间布局,生成多样化新全景图像以丰富数据集。我们首先将每张全景图像分解为组成要素:背景风格、前景家具和房间布局。然后,通过混合三张不同图像中的这三个部分(例如一张图像的前景家具、另一张图像的背景风格、第三张图像的房间结构)来生成增强图像。由于图像组合呈立方级增长,该方法具有高度多样性。我们还评估了PanoMixSwap在语义分割和布局估计这两项室内场景理解任务上的有效性。实验表明,使用PanoMixSwap训练的最先进方法在这两项任务上均持续优于原始设置。