Diffusion models (DMs) excel in photo-realistic image synthesis, but their adaptation to LiDAR scene generation poses a substantial hurdle. This is primarily because DMs operating in the point space struggle to preserve the curve-like patterns and 3D geometry of LiDAR scenes, which consumes much of their representation power. In this paper, we propose LiDAR Diffusion Models (LiDMs) to generate LiDAR-realistic scenes from a latent space tailored to capture the realism of LiDAR scenes by incorporating geometric priors into the learning pipeline. Our method targets three major desiderata: pattern realism, geometry realism, and object realism. Specifically, we introduce curve-wise compression to simulate real-world LiDAR patterns, point-wise coordinate supervision to learn scene geometry, and patch-wise encoding for a full 3D object context. With these three core designs, our method achieves competitive performance on unconditional LiDAR generation in 64-beam scenario and state of the art on conditional LiDAR generation, while maintaining high efficiency compared to point-based DMs (up to 107$\times$ faster). Furthermore, by compressing LiDAR scenes into a latent space, we enable the controllability of DMs with various conditions such as semantic maps, camera views, and text prompts.
翻译:扩散模型(DMs)在照片级真实感图像合成方面表现出色,但其在激光雷达场景生成中的应用面临重大挑战。这主要是因为工作在点空间的扩散模型难以保持激光雷达场景中的曲线状模式和三维几何结构,消耗了大量表征能力。本文提出激光雷达扩散模型(LiDMs),通过将几何先验融入学习流程,从专为捕捉激光雷达场景真实感而设计的潜在空间生成逼真的激光雷达场景。我们的方法针对三个核心目标:模式真实性、几何真实性和物体真实性。具体而言,我们引入曲线级压缩以模拟真实激光雷达模式,采用点级坐标监督学习场景几何结构,并利用补丁级编码实现完整的3D物体上下文。通过这三项核心设计,我们的方法在64线束场景的无条件激光雷达生成中取得了竞争性性能,在有条件激光雷达生成中达到最先进水平,同时相比基于点的扩散模型保持高效率(速度提升高达107倍)。此外,通过将激光雷达场景压缩至潜在空间,我们实现了扩散模型在语义地图、相机视图和文本提示等多种条件下的可控性。