This paper presents a novel approach, TeFS (Temporal-controlled Frame Swap), to generate synthetic stereo driving data for visual simultaneous localization and mapping (vSLAM) tasks. TeFS is designed to overcome the lack of native stereo vision support in commercial driving simulators, and we demonstrate its effectiveness using Grand Theft Auto V (GTA V), a high-budget open-world video game engine. We introduce GTAV-TeFS, the first large-scale GTA V stereo-driving dataset, containing over 88,000 high-resolution stereo RGB image pairs, along with temporal information, GPS coordinates, camera poses, and full-resolution dense depth maps. GTAV-TeFS offers several advantages over other synthetic stereo datasets and enables the evaluation and enhancement of state-of-the-art stereo vSLAM models under GTA V's environment. We validate the quality of the stereo data collected using TeFS by conducting a comparative analysis with the conventional dual-viewport data using an open-source simulator. We also benchmark various vSLAM models using the challenging-case comparison groups included in GTAV-TeFS, revealing the distinct advantages and limitations inherent to each model. The goal of our work is to bring more high-fidelity stereo data from commercial-grade game simulators into the research domain and push the boundary of vSLAM models. %Our dataset also demonstrates the effectiveness of pre-trained state-of-the-art stereo matching networks, which show considerable performance gains on KITTI stereo depth estimation benchmarks. All code and datasets will be released upon acceptance.
翻译:本文提出了一种名为TeFS(时间控制帧交换)的新方法,用于生成用于视觉同步定位与地图构建(vSLAM)任务的合成立体驾驶数据。TeFS旨在解决商业驾驶模拟器缺乏原生立体视觉支持的问题,并通过使用高预算开放世界视频游戏引擎《侠盗猎车手V》(GTA V)证明了其有效性。我们推出了GTAV-TeFS,这是首个大规模GTA V立体驾驶数据集,包含超过88,000对高分辨率立体RGB图像对,以及时间信息、GPS坐标、相机位姿和全分辨率密集深度图。GTAV-TeFS相比其他合成立体数据集具有多项优势,能够在GTA V环境下评估和增强最先进的立体vSLAM模型。我们通过使用开源模拟器,将TeFS采集的立体数据与传统的双视口数据进行对比分析,验证了所采集立体数据的质量。我们还利用GTAV-TeFS中包含的挑战性案例对比组,对各种vSLAM模型进行了基准测试,揭示了每个模型固有的独特优势和局限性。本研究的目标是将更多来自商业级游戏模拟器的高保真立体数据引入研究领域,并推动vSLAM模型的发展。%我们的数据集还展示了预训练最先进立体匹配网络的有效性,这些网络在KITTI立体深度估计基准上取得了显著的性能提升。所有代码和数据集将在接收后公开发布。