Multimodal perception systems for robotics and embodied AI often assume reliable RGB-D sensing, but in practice, depth is frequently missing, noisy, or corrupted. We thus present GeomPrompt, a lightweight cross-modal adaptation module that synthesizes a task-driven geometric prompt from RGB alone for the fourth channel of a frozen RGB-D semantic segmentation model, without depth supervision. We further introduce GeomPrompt-Recovery, an adaptation module that compensates for degraded depth by predicting the fourth channel correction relevant for the frozen segmenter. Both modules are trained solely with downstream segmentation supervision, enabling recovery of the geometric prior useful for segmentation, rather than estimating depth signals. On SUN RGB-D, GeomPrompt improves over RGB-only inference by +6.1 mIoU on DFormer and +3.0 mIoU on GeminiFusion, while remaining competitive with strong monocular depth estimators. For degraded depth, GeomPrompt-Recovery consistently improves robustness, yielding gains up to +3.6 mIoU under severe depth corruptions. GeomPrompt is also substantially more efficient than monocular depth baselines, reaching 7.8 ms latency versus 38.3 ms and 71.9 ms. These results suggest that task-driven geometric prompting is an efficient mechanism for cross-modal compensation under missing and degraded depth inputs in RGB-D perception.
翻译:摘要:面向机器人和具身智能的多模态感知系统常假设RGB-D传感可靠,但实际中深度信息频繁缺失、含噪或损坏。为此,我们提出GeomPrompt——一种轻量级跨模态适配模块,仅从RGB图像合成任务驱动的几何提示,作为冻结RGB-D语义分割模型的第四通道输入,无需深度监督。进一步提出GeomPrompt-Recovery适配模块,通过预测冻结分割器所需的第四通道校正量来补偿退化深度。两个模块仅依赖下游分割监督训练,旨在恢复对分割有用的几何先验而非直接估计深度信号。在SUN RGB-D数据集上,GeomPrompt相比纯RGB推理在DFormer上提升+6.1 mIoU,在GeminiFusion上提升+3.0 mIoU,且性能与强单目深度估计器相当。针对退化深度场景,GeomPrompt-Recovery持续提升鲁棒性,在严重深度损坏条件下取得高达+3.6 mIoU的性能增益。GeomPrompt在效率上显著优于单目深度基线,延迟仅7.8毫秒,对比38.3毫秒和71.9毫秒。结果表明,任务驱动的几何提示是RGB-D感知中应对缺失与退化深度输入的高效跨模态补偿机制。