Although LiDAR semantic segmentation advances rapidly, state-of-the-art methods often incorporate specifically designed inductive bias derived from benchmarks originating from mechanical spinning LiDAR. This can limit model generalizability to other kinds of LiDAR technologies and make hyperparameter tuning more complex. To tackle these issues, we propose a generalized framework to accommodate various types of LiDAR prevalent in the market by replacing window-attention with our sparse focal point modulation. Our SFPNet is capable of extracting multi-level contexts and dynamically aggregating them using a gate mechanism. By implementing a channel-wise information query, features that incorporate both local and global contexts are encoded. We also introduce a novel large-scale hybrid-solid LiDAR semantic segmentation dataset for robotic applications. SFPNet demonstrates competitive performance on conventional benchmarks derived from mechanical spinning LiDAR, while achieving state-of-the-art results on benchmark derived from solid-state LiDAR. Additionally, it outperforms existing methods on our novel dataset sourced from hybrid-solid LiDAR. Code and dataset are available at https://github.com/Cavendish518/SFPNet and https://www.semanticindustry.top.
翻译:尽管激光雷达语义分割技术发展迅速,但当前最先进的方法通常融合了源自机械旋转式激光雷达基准测试的特定归纳偏置。这可能会限制模型对其他类型激光雷达技术的泛化能力,并增加超参数调优的复杂性。为解决这些问题,我们提出了一种通用框架,通过用我们提出的稀疏焦点调制机制替代窗口注意力机制,以适应市场上主流的各类激光雷达。我们的SFPNet能够提取多层级上下文信息,并利用门控机制进行动态聚合。通过实施通道级信息查询,编码了融合局部与全局上下文的特征。我们还为机器人应用引入了一个新颖的大规模混合固态激光雷达语义分割数据集。SFPNet在基于机械旋转式激光雷达的传统基准测试中展现出具有竞争力的性能,同时在基于固态激光雷达的基准测试中取得了最先进的结果。此外,在我们提出的混合固态激光雷达新数据集上,其性能也超越了现有方法。代码与数据集可通过https://github.com/Cavendish518/SFPNet 与 https://www.semanticindustry.top 获取。