Prior methods that tackle the problem of generalizable object pose estimation highly rely on having dense views of the unseen object. By contrast, we address the scenario where only a single reference view of the object is available. Our goal then is to estimate the relative object pose between this reference view and a query image that depicts the object in a different pose. In this scenario, robust generalization is imperative due to the presence of unseen objects during testing and the large-scale object pose variation between the reference and the query. To this end, we present a new hypothesis-and-verification framework, in which we generate and evaluate multiple pose hypotheses, ultimately selecting the most reliable one as the relative object pose. To measure reliability, we introduce a 3D-aware verification that explicitly applies 3D transformations to the 3D object representations learned from the two input images. Our comprehensive experiments on the Objaverse, LINEMOD, and CO3D datasets evidence the superior accuracy of our approach in relative pose estimation and its robustness in large-scale pose variations, when dealing with unseen objects.
翻译:现有方法在处理泛化物体姿态估计问题时,高度依赖于对未见物体密集视角的获取。相比之下,我们针对仅存在物体单一参考视图的场景展开研究。目标是在参考视图与描述同一物体不同姿态的查询图像之间,估计其相对物体姿态。在此场景中,由于测试阶段存在未见物体,且参考视图与查询图像之间存在大规模姿态变化,鲁棒泛化能力至关重要。为此,我们提出一种新的假设验证框架,通过生成并评估多个姿态假设,最终选取最可靠的相对物体姿态。在可靠性评估方面,我们引入三维感知验证机制,该机制显式地将三维变换应用于从两幅输入图像学习到的三维物体表征。我们在Objaverse、LINEMOD和CO3D数据集上的全面实验表明,本方法在应对未见物体时,在相对姿态估计精度及大规模姿态变化鲁棒性方面均展现出显著优势。