Vision research showed remarkable success in understanding our world, propelled by datasets of images and videos. Sensor data from radar, LiDAR and cameras supports research in robotics and autonomous driving for at least a decade. However, while visual sensors may fail in some conditions, sound has recently shown potential to complement sensor data. Simulated room impulse responses (RIR) in 3D apartment-models became a benchmark dataset for the community, fostering a range of audiovisual research. In simulation, depth is predictable from sound, by learning bat-like perception with a neural network. Concurrently, the same was achieved in reality by using RGB-D images and echoes of chirping sounds. Biomimicking bat perception is an exciting new direction but needs dedicated datasets to explore the potential. Therefore, we collected the BatVision dataset to provide large-scale echoes in complex real-world scenes to the community. We equipped a robot with a speaker to emit chirps and a binaural microphone to record their echoes. Synchronized RGB-D images from the same perspective provide visual labels of traversed spaces. We sampled modern US office spaces to historic French university grounds, indoor and outdoor with large architectural variety. This dataset will allow research on robot echolocation, general audio-visual tasks and sound ph{\ae}nomena unavailable in simulated data. We show promising results for audio-only depth prediction and show how state-of-the-art work developed for simulated data can also succeed on our dataset. Project page: https://amandinebtto.github.io/Batvision-Dataset/
翻译:视觉研究在理解世界方面取得了显著成功,这得益于图像和视频数据集的发展。雷达、激光雷达和摄像头等传感器数据支持机器人和自动驾驶研究已有十余年之久。然而,尽管视觉传感器在某些条件下可能失效,但声音近年来展现出补充传感器数据的潜力。基于三维公寓模型的模拟房间脉冲响应(RIR)已成为该领域的基准数据集,推动了多种视听研究。在模拟环境中,通过使用神经网络学习类似蝙蝠的感知能力,可以从声音中预测深度。与此同时,在现实场景中,通过使用RGB-D图像和啁啾声回声也实现了同样的效果。仿生蝙蝠感知是一个令人兴奋的新方向,但需要专门的数据集来探索其潜力。因此,我们收集了BatVision数据集,为复杂真实世界场景中的大规模回声研究提供支持。我们为一台机器人配备了扬声器以发出啁啾声,以及双耳麦克风以记录其回声。同一视角下的同步RGB-D图像提供了所经过空间的视觉标签。我们采样了从现代美国办公空间到历史悠久的法国大学校园、室内外兼具多样建筑风格的场景。该数据集将支持机器人回声定位、通用视听任务以及模拟数据中无法获得的声学现象研究。我们展示了音频仅深度预测的初步结果,并证明了为模拟数据开发的最先进方法也能在我们的数据集上成功应用。项目页面:https://amandinebtto.github.io/Batvision-Dataset/