Crowd-sourced cooperative mapping from monocular cameras promises scalable 3D reconstruction without specialized sensors, yet remains hindered by two scale-specific failure modes: abrupt scale collapse from false-positive loop closures in repetitive environments, and gradual scale drift over long trajectories and per-robot scale ambiguity that prevent direct multi-session fusion. We present MR.ScaleMaster, a cooperative mapping system for crowd-sourced monocular videos that addresses both failure modes. MR.ScaleMaster introduces three key mechanisms. First, a Scale Collapse Alarm rejects spurious loop closures before they corrupt the pose graph. Second, a Sim(3) anchor node formulation generalizes the classical SE(3) framework to explicitly estimate per-session scale, resolving per-robot scale ambiguity and enforcing global scale consistency. Third, a modular, open-source, plug-and-play interface enables any monocular reconstruction model to integrate without backend modification. On KITTI sequences with up to 15 agents, the Sim(3) formulation achieves a 7.2x ATE reduction over the SE(3) baseline, and the alarm rejects all false-positive loops while preserving every valid constraint. We further demonstrate heterogeneous multi-robot dense mapping fusing MASt3R-SLAM, pi3, and VGGT-SLAM 2.0 within a single unified map.
翻译:众包单目相机的协同建图旨在无需专用传感器即可实现可扩展的三维重建,然而其发展受到两种特定尺度故障模式的阻碍:在重复环境中由假阳性回环检测导致的突然尺度崩溃,以及在长轨迹上因单机器人尺度模糊性引发且阻碍多会话直接融合的渐进尺度漂移。我们提出MR.ScaleMaster——一种针对众包单目视频的协同建图系统,可同时应对上述两种故障模式。该系统引入三项关键机制:第一,尺度崩溃警报机制在虚假回环破坏位姿图之前将其剔除;第二,Sim(3)锚点节点公式将经典SE(3)框架推广至显式估计每会话尺度,消除单机器人尺度模糊性并强制执行全局尺度一致性;第三,模块化、开源且即插即用的接口使任意单目重建模型无需修改后端即可集成。在包含最多15个智能体的KITTI序列上,Sim(3)公式相较SE(3)基线实现了7.2倍的绝对轨迹误差降低,警报机制在保留所有有效约束的同时剔除了所有假阳性回环。我们进一步演示了融合MASt3R-SLAM、pi3与VGGT-SLAM 2.0的异构多机器人稠密建图,在单一统一地图中实现协同。