Estimating depth from images nowadays yields outstanding results, both in terms of in-domain accuracy and generalization. However, we identify two main challenges that remain open in this field: dealing with non-Lambertian materials and effectively processing high-resolution images. Purposely, we propose a novel dataset that includes accurate and dense ground-truth labels at high resolution, featuring scenes containing several specular and transparent surfaces. Our acquisition pipeline leverages a novel deep space-time stereo framework, enabling easy and accurate labeling with sub-pixel precision. The dataset is composed of 606 samples collected in 85 different scenes, each sample includes both a high-resolution pair (12 Mpx) as well as an unbalanced stereo pair (Left: 12 Mpx, Right: 1.1 Mpx), typical of modern mobile devices that mount sensors with different resolutions. Additionally, we provide manually annotated material segmentation masks and 15K unlabeled samples. The dataset is composed of a train set and two test sets, the latter devoted to the evaluation of stereo and monocular depth estimation networks. Our experiments highlight the open challenges and future research directions in this field.
翻译:当前,基于图像的深度估计方法在域内精度和泛化能力方面均取得了显著成果。然而,我们识别出该领域仍存在的两大核心挑战:处理非朗伯体材质以及高效解析高分辨率图像。为此,我们提出一个包含高分辨率、高精度密集真值标签的新型数据集,其场景中包含了多种镜面与透明表面。我们采用创新的深度时空立体框架构建采集管线,可实现亚像素精度的简易且准确标注。该数据集收录了85个不同场景的606组样本,每组样本同时包含高分辨率图像对(12百万像素)与非均衡立体图像对(左12百万像素/右1.1百万像素),后者典型表征了现代移动设备采用多分辨率传感器的特性。此外,我们提供了人工标注的材质分割掩码和15,000张无标注样本。数据集包含训练集与两个测试集,后者分别用于评估立体与单目深度估计网络性能。实验揭示了该领域当前面临的开放性挑战与未来研究方向。