We introduce a new diffusion-based approach for shape completion on 3D range scans. Compared with prior deterministic and probabilistic methods, we strike a balance between realism, multi-modality, and high fidelity. We propose DiffComplete by casting shape completion as a generative task conditioned on the incomplete shape. Our key designs are two-fold. First, we devise a hierarchical feature aggregation mechanism to inject conditional features in a spatially-consistent manner. So, we can capture both local details and broader contexts of the conditional inputs to control the shape completion. Second, we propose an occupancy-aware fusion strategy in our model to enable the completion of multiple partial shapes and introduce higher flexibility on the input conditions. DiffComplete sets a new SOTA performance (e.g., 40% decrease on l_1 error) on two large-scale 3D shape completion benchmarks. Our completed shapes not only have a realistic outlook compared with the deterministic methods but also exhibit high similarity to the ground truths compared with the probabilistic alternatives. Further, DiffComplete has strong generalizability on objects of entirely unseen classes for both synthetic and real data, eliminating the need for model re-training in various applications.
翻译:我们提出了一种基于扩散的新方法,用于三维距离扫描的形状补全。与先前的确定性和概率性方法相比,我们在真实性、多模态和高保真度之间取得了平衡。通过将形状补全视为以不完整形状为条件的生成任务,我们提出了DiffComplete。我们的关键设计体现在两方面。首先,我们设计了一种层级特征聚合机制,以空间一致的方式注入条件特征,从而能够捕捉条件输入的局部细节和更广泛的上下文以控制形状补全。其次,我们在模型中提出了一种基于占用率的融合策略,能够完成多个部分形状的补全,并为输入条件引入更高的灵活性。DiffComplete在两个大规模三维形状补全基准上取得了新的最优性能(例如,l₁误差降低了40%)。与确定性方法相比,我们补全的形状不仅具有逼真的外观,而且与概率性方法相比,展现出与真实值的高度相似性。此外,DiffComplete对合成数据和真实数据中完全未见类别的物体均表现出强大的泛化能力,消除了各种应用中模型重新训练的需求。