LiDAR sensors play a crucial role in various applications, especially in autonomous driving. Current research primarily focuses on optimizing perceptual models with point cloud data as input, while the exploration of deeper cognitive intelligence remains relatively limited. To address this challenge, parallel LiDARs have emerged as a novel theoretical framework for the next-generation intelligent LiDAR systems, which tightly integrate physical, digital, and social systems. To endow LiDAR systems with cognitive capabilities, we introduce the 3D visual grounding task into parallel LiDARs and present a novel human-computer interaction paradigm for LiDAR systems. We propose Talk2LiDAR, a large-scale benchmark dataset tailored for 3D visual grounding in autonomous driving. Additionally, we present a two-stage baseline approach and an efficient one-stage method named BEVGrounding, which significantly improves grounding accuracy by fusing coarse-grained sentence and fine-grained word embeddings with visual features. Our experiments on Talk2Car-3D and Talk2LiDAR datasets demonstrate the superior performance of BEVGrounding, laying a foundation for further research in this domain.
翻译:激光雷达传感器在各类应用中,尤其在自动驾驶领域,发挥着至关重要的作用。当前研究主要集中于以点云数据为输入优化感知模型,而对更深层次认知智能的探索仍相对有限。为应对这一挑战,平行激光雷达作为一种新颖的理论框架应运而生,旨在构建下一代智能激光雷达系统,其紧密融合了物理、数字与社会系统。为使激光雷达系统具备认知能力,我们将三维视觉定位任务引入平行激光雷达,并提出了一种新颖的激光雷达系统人机交互范式。我们提出了Talk2LiDAR,一个专为自动驾驶场景下三维视觉定位定制的大规模基准数据集。此外,我们提出了一种两阶段基线方法以及一种高效的单阶段方法BEVGrounding,该方法通过将粗粒度句子嵌入与细粒度词嵌入同视觉特征相融合,显著提升了定位精度。我们在Talk2Car-3D和Talk2LiDAR数据集上的实验证明了BEVGrounding的优越性能,为该领域的进一步研究奠定了基础。