Reliable monocular video depth estimation is crucial for downstream 3D reasoning and embodied AI in endoscopic navigation. However, existing self-supervised approaches typically treat video frames independently or rely on weak temporal regularization. These methods, lacking a holistic perception of the underlying 3D scene, inevitably suffer from geometrically inconsistent predictions and severe cross-frame drift. To address these limitations, we introduce a new paradigm that recasts sequential video depth estimation as an unconstrained multi-view 3D reconstruction problem, enabling full exploitation of the powerful geometric priors embedded in recent 3D foundation models. The core of our approach is a 3D consistency optimization framework driven by three constraints: image-level photometric rendering, explicit world-coordinate geometric alignment, and multi-scale temporal gradient consistency. Such unified optimization elegantly anchors isolated frames to a globally coherent 3D structure. Our method has been validated in both the self-supervised training scenarios and challenging zero-shot clinical environments. Results show that the proposed approach achieves state-of-the-art spatial accuracy, outperforming the frame-based, video-based depth estimators and the multi-view 3D reconstruction baselines.


翻译:可靠的单目视频深度估计对于下游3D推理和腔内导航中的具身AI至关重要。然而,现有自监督方法通常独立处理视频帧或依赖弱时间正则化。这些方法缺乏对底层3D场景的整体感知,不可避免地导致几何不一致的预测和严重的跨帧漂移。为解决这些局限,我们提出一种新范式,将顺序视频深度估计重新定义为无约束的多视角3D重建问题,从而充分利用嵌入在最新3D基础模型中的强大几何先验。我们方法的核心是一个由三种约束驱动的3D一致性优化框架:图像级光度渲染、显式世界坐标几何对齐和多尺度时间梯度一致性。这种统一优化优雅地将孤立帧锚定到全局一致的3D结构上。我们的方法已在自监督训练场景和具有挑战性的零样本临床环境中得到验证。结果表明,所提方法实现了最先进的空间精度,优于基于帧和视频的深度估计器以及多视角3D重建基线。

0
下载
关闭预览

相关内容

3D是英文“Three Dimensions”的简称,中文是指三维、三个维度、三个坐标,即有长、有宽、有高,换句话说,就是立体的,是相对于只有长和宽的平面(2D)而言。
迈向深度基础模型:基于视觉的深度估计最新趋势
专知会员服务
24+阅读 · 2025年7月16日
【博士论文】基于深度学习的单目场景深度估计方法研究
MonoGRNet:单目3D目标检测的通用框架(TPAMI2021)
专知会员服务
18+阅读 · 2021年5月3日
专知会员服务
65+阅读 · 2021年4月11日
半监督深度学习小结:类协同训练和一致性正则化
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
9+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
9+阅读 · 8月1日
相关VIP内容
迈向深度基础模型:基于视觉的深度估计最新趋势
专知会员服务
24+阅读 · 2025年7月16日
【博士论文】基于深度学习的单目场景深度估计方法研究
MonoGRNet:单目3D目标检测的通用框架(TPAMI2021)
专知会员服务
18+阅读 · 2021年5月3日
专知会员服务
65+阅读 · 2021年4月11日
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员