Reconstructing hand-held objects from a single RGB image without known 3D object templates, category prior, or depth information is a vital yet challenging problem in computer vision. In contrast to prior works that utilize deterministic modeling paradigms, which make it hard to account for the uncertainties introduced by hand- and self-occlusion, we employ a probabilistic point cloud denoising diffusion model to tackle the above challenge. In this work, we present Hand-Aware Conditional Diffusion for monocular hand-held object reconstruction (HACD), modeling the hand-object interaction in two aspects. First, we introduce hand-aware conditioning to model hand-object interaction from both semantic and geometric perspectives. Specifically, a unified hand-object semantic embedding compensates for the 2D local feature deficiency induced by hand occlusion, and a hand articulation embedding further encodes the relationship between object vertices and hand joints. Second, we propose a hand-constrained centroid fixing scheme, which utilizes hand vertices priors to restrict the centroid deviation of partially denoised point cloud during diffusion and reverse process. Removing the centroid bias interference allows the diffusion models to focus on the reconstruction of shape, thus enhancing the stability and precision of local feature projection. Experiments on the synthetic ObMan dataset and two real-world datasets, HO3D and MOW, demonstrate our approach surpasses all existing methods by a large margin.
翻译:从单个RGB图像中重建手持物体,且无需已知3D物体模板、类别先验或深度信息,是计算机视觉领域重要且具有挑战性的问题。不同于现有采用确定性建模范式(难以应对手部自遮挡引发的不确定性)的研究,我们采用概率性点云去噪扩散模型应对上述挑战。本文提出面向单目手持物体重建的手部感知条件扩散模型(HACD),从两方面建模手-物交互:首先,引入手部感知条件机制,从语义与几何双重维度建模手物交互——统一的手物语义嵌入弥补手部遮挡导致的二维局部特征缺失,手部关节嵌入进一步编码物体顶点与手部关节的关联;其次,提出手部约束质心固定方案,利用手部顶点先验限制扩散与逆过程中部分去噪点云的质心偏移。消除质心偏差干扰使扩散模型专注于形状重建,从而增强局部特征投影的稳定性与精度。在合成数据集ObMan及两个真实数据集HO3D、MOW上的实验表明,本方法以显著优势超越所有现有方法。