While monocular depth estimation has achieved significant progress, achieving generalized metric depth estimation for both narrow field-of-view (FoV) perspectives and $360^\circ$ panoramas remains an unsolved challenge. Existing methods are often tailored to specific camera types and struggle to produce accurate metric depth that generalizes across diverse settings. This limitation stems from two key challenges: the inherent geometric discrepancy between perspective and panoramic cameras, and the scarcity of panoramic training data with metric annotations. In this work, we introduce DepthMaster, a unified metric depth estimation framework. Rather than employing specialized networks to learn spherical distortions, we reformulate the problem by decomposing panoramic images into overlapping perspective patches. Crucially, distinct from prior projection-based methods that rely on ad-hoc architectural modifications to handle boundaries, we introduce a novel Correspondence Consistency Loss (CCL) and inject virtual projection cameras as geometric priors, allowing us to seamlessly stitch the patches while avoiding specialized operators and keeping the backbone largely compatible with standard Transformer designs. This strategy also resolves the geometric differences by unifying all inputs into a canonical perspective representation, and effectively circumvents data scarcity by directly unlocking powerful metric priors from vast perspective datasets. Trained on a mixed dataset that contains only one panorama dataset, DepthMaster achieves state-of-the-art zero-shot performance on 13 diverse datasets, outperforming not only universal methods but also leading specialist models in both perspective and panoramic domains.


翻译:[translated abstract in Chinese] 尽管单目深度估计已取得显著进展,但针对窄视场角(FoV)透视图像与$360^\circ$全景图像实现泛化性公制深度估计仍是一项未解决的挑战。现有方法往往针对特定相机类型量身定制,难以在跨多样化场景设置下生成泛化的准确公制深度。这一局限性源于两个关键难题:透视相机与全景相机之间固有的几何差异,以及带有公制标注的全景训练数据的匮乏。在本工作中,我们提出DepthMaster——一种统一的公制深度估计框架。该方法并未采用专用网络学习球面畸变,而是通过将全景图像分解为重叠透视补丁来重新定义问题。关键在于,不同于以往依赖专门架构修改以处理边界的投影类方法,我们引入了一种新颖的对应一致性损失(CCL),并注入虚拟投影相机作为几何先验,从而在无需专用算子的前提下实现补丁的无缝拼接,同时保持主干网络与标准Transformer设计的广泛兼容性。该策略通过将所有输入统一至标准透视表示来解决几何差异,并通过直接解锁大规模透视数据集中强大的公制先验,有效规避了数据稀缺问题。在仅包含一个全景数据集的混合数据集上训练后,DepthMaster在13个多样化数据集上实现了最先进的零样本性能,不仅超越通用方法,还在透视与全景两个领域领先于专业模型。

0
下载
关闭预览

相关内容

迈向深度基础模型:基于视觉的深度估计最新趋势
专知会员服务
24+阅读 · 2025年7月16日
【博士论文】基于深度学习的单目场景深度估计方法研究
专知会员服务
42+阅读 · 2021年9月30日
专知会员服务
35+阅读 · 2021年2月7日
【旷视出品】细粒度图像分析综述
专知
15+阅读 · 2019年7月11日
深度相机原理揭秘--双目立体视觉
计算机视觉life
10+阅读 · 2017年11月7日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
致命七类无人机:无人机时代的演进型合成兵种
《异构无人水面艇集群作战自主制导算法》130页
相关VIP内容
迈向深度基础模型:基于视觉的深度估计最新趋势
专知会员服务
24+阅读 · 2025年7月16日
【博士论文】基于深度学习的单目场景深度估计方法研究
专知会员服务
42+阅读 · 2021年9月30日
专知会员服务
35+阅读 · 2021年2月7日
相关基金
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员