In recent years, there has been a growing trend of using data-driven methods in industrial settings. These kinds of methods often process video images or parts, therefore the integrity of such images is crucial. Sometimes datasets, e.g. consisting of images, can be sophisticated for various reasons. It becomes critical to understand how the manipulation of video and images can impact the effectiveness of a machine learning method. Our case study aims precisely to analyze the Linemod dataset, considered the state of the art in 6D pose estimation context. That dataset presents images accompanied by ArUco markers; it is evident that such markers will not be available in real-world contexts. We analyze how the presence of the markers affects the pose estimation accuracy, and how this bias may be mitigated through data augmentation and other methods. Our work aims to show how the presence of these markers goes to modify, in the testing phase, the effectiveness of the deep learning method used. In particular, we will demonstrate, through the tool of saliency maps, how the focus of the neural network is captured in part by these ArUco markers. Finally, a new dataset, obtained by applying geometric tools to Linemod, will be proposed in order to demonstrate our hypothesis and uncovering the bias. Our results demonstrate the potential for bias in 6DOF pose estimation networks, and suggest methods for reducing this bias when training with markers.
翻译:近年来,在工业场景中应用数据驱动方法的趋势日益增长。此类方法通常处理视频图像或部件,因此图像的完整性至关重要。由于多种原因,数据集(例如由图像组成的数据集)可能变得复杂。理解视频和图像的操作如何影响机器学习方法的有效性至关重要。我们的案例研究旨在精确分析Linemod数据集——该数据集被视为6D姿态估计领域的当前最优基准。该数据集中的图像包含ArUco标记,而在实际场景中显然不存在此类标记。我们分析了标记存在对姿态估计精度的影响,以及如何通过数据增强等方法减轻这种偏差。我们的工作旨在展示这些标记如何在测试阶段改变所用深度学习方法的有效性。具体而言,我们将通过显著图工具证明,神经网络的注意力部分被这些ArUco标记所吸引。最后,我们提出一个基于Linemod通过几何工具处理的新数据集,以验证我们的假设并揭示偏差。实验结果展示了6DOF姿态估计网络中潜在的偏差,并提出了在使用标记进行训练时减少此类偏差的方法。