Autonomous robots operating in natural karstic caves face perception and navigation challenges that are qualitatively distinct from those encountered in mines or tunnels: irregular geometry, reflective wet surfaces, near-zero ambient light, and complex branching passages. Yet publicly available datasets targeting this environment remain scarce and offer limited sensing modalities and environmental diversity. We present CAVERS, a multimodal dataset acquired in two structurally distinct rooms of Cueva de la Victoria, Málaga, Spain, comprising 24 sequences totaling approximately 335 GB of recorded data. The sensor suite combines an Intel RealSense D435i RGB-D-I camera, an Optris PI640i near-IR thermal camera, and a Velodyne VLP-16 LiDAR, operated both handheld and mounted on a wheeled rover under full darkness and artificial illumination. For most of the sequences, mm-accurate 6-DoF ground truth pose and velocity at 120 Hz are provided by an Optirack motion capture system installed directly inside the cave. We benchmark seven state-of-the-art SLAM and odometry algorithms spanning visual, visual-inertial, thermal-inertial, and LiDAR-based pipelines, as well as a 3D reconstruction pipeline, demonstrating the dataset's usability. %The dataset and all supplementary material are publicly available at: https://github.com/spaceuma/cavers.
翻译:在天然溶洞中运行的自主机器人面临着与矿井或隧道截然不同的感知与导航挑战:不规则几何结构、反射性潮湿表面、接近零的环境光照以及复杂的分支通道。然而,针对这一环境的公开数据集仍然稀少,且传感模态和场景多样性有限。我们提出CAVERS,这是一个在西班牙马拉加维多利亚洞穴中两处结构差异显著的空间内采集的多模态数据集,包含24个序列,总计约335 GB的记录数据。传感器套件结合了Intel RealSense D435i RGB-D-I相机、Optris PI640i近红外热成像相机以及Velodyne VLP-16激光雷达,在完全黑暗和人工照明的条件下,分别以手持和轮式漫游车搭载的方式运行。对于大多数序列,由直接安装在洞穴内的Optirack运动捕捉系统提供120 Hz频率的毫米级精确6自由度位姿与速度真值。我们评估了七种涵盖视觉、视觉-惯性、热成像-惯性及基于激光雷达管道的最先进SLAM与里程计算法,以及一种三维重建管道,验证了该数据集的实用性。%数据集及所有补充材料公开获取地址:https://github.com/spaceuma/cavers。