This paper addresses the challenging problem of category-level pose estimation. Current state-of-the-art methods for this task face challenges when dealing with symmetric objects and when attempting to generalize to new environments solely through synthetic data training. In this work, we address these challenges by proposing a probabilistic model that relies on diffusion to estimate dense canonical maps crucial for recovering partial object shapes as well as establishing correspondences essential for pose estimation. Furthermore, we introduce critical components to enhance performance by leveraging the strength of the diffusion models with multi-modal input representations. We demonstrate the effectiveness of our method by testing it on a range of real datasets. Despite being trained solely on our generated synthetic data, our approach achieves state-of-the-art performance and unprecedented generalization qualities, outperforming baselines, even those specifically trained on the target domain.
翻译:本文针对类别级姿态估计这一挑战性问题展开研究。现有最先进方法在处理对称物体以及仅通过合成数据训练泛化至新环境时仍面临困难。为此,我们提出了一种基于扩散的概率模型,用于估计密集规范映射——该映射对于恢复部分物体形状及建立姿态估计所需的关键对应关系至关重要。此外,我们引入关键组件,通过利用扩散模型在多模态输入表征中的优势来提升性能。通过在多个真实数据集上进行测试,我们验证了本方法的有效性。尽管仅使用自生成的合成数据进行训练,本方法仍实现了最先进的性能与前所未有的泛化能力,显著优于基线方法——甚至包括那些在目标域上专门训练的模型。