Recent Multimodal Large Language Models (MLLMs) have shown high potential for spatial reasoning within 3D scenes. However, they typically rely on computationally expensive 3D representations like point clouds or reconstructed Bird's-Eye View (BEV) maps, or lack physical grounding to resolve ambiguities in scale and size. This paper significantly enhances MLLMs with egomotion modality data, captured by Inertial Measurement Units (IMUs) concurrently with the video. In particular, we propose a novel framework, called Motion-MLLM, introducing two key components: (1) a cascaded motion-visual keyframe filtering module that leverages both IMU data and visual features to efficiently select a sparse yet representative set of keyframes, and (2) an asymmetric cross-modal fusion module where motion tokens serve as intermediaries that channel egomotion cues and cross-frame visual context into the visual representation. By grounding visual content in physical egomotion trajectories, Motion-MLLM can reason about absolute scale and spatial relationships across the scene. Our extensive evaluation shows that Motion-MLLM makes significant improvements in various tasks related to 3D scene understanding and spatial reasoning. Compared to state-of-the-art (SOTA) methods based on video frames and explicit 3D data, Motion-MLLM exhibits similar or even higher accuracy with significantly less overhead (i.e., 1.40$\times$ and 1.63$\times$ higher cost-effectiveness, respectively).
翻译:近期多模态大语言模型(MLLMs)在3D场景内的空间推理方面展现出巨大潜力。然而,它们通常依赖计算成本高昂的3D表征(如点云或重建的鸟瞰图(BEV)),或缺乏物理基础以解决尺度与尺寸歧义。本文通过自运动模态数据(由惯性测量单元(IMUs)与视频同步采集)显著增强MLLMs。具体而言,我们提出名为Motion-MLLM的新框架,引入两个关键组件:(1)级联运动-视觉关键帧筛选模块,利用IMU数据与视觉特征高效选取稀疏且具代表性的关键帧集合;(2)非对称跨模态融合模块,其中运动令牌作为媒介,将自运动线索与跨帧视觉上下文注入视觉表征。通过将视觉内容锚定于物理自运动轨迹,Motion-MLLM能够推理场景中的绝对尺度与空间关系。广泛评估表明,Motion-MLLM在多项与3D场景理解及空间推理相关的任务中取得显著提升。与基于视频帧与显式3D数据的先进方法(SOTA)相比,Motion-MLLM在显著降低开销的同时展现了相当或更高的精度(成本效益分别提升1.40倍与1.63倍)。