Category-agnostic pose estimation (CAPE) aims to predict keypoints for arbitrary classes given a few support images annotated with keypoints. Existing methods only rely on the features extracted at support keypoints to predict or refine the keypoints on query image, but a few support feature vectors are local and inadequate for CAPE. Considering that human can quickly perceive potential keypoints of arbitrary objects, we propose a novel framework for CAPE based on such potential keypoints (named as meta-points). Specifically, we maintain learnable embeddings to capture inherent information of various keypoints, which interact with image feature maps to produce meta-points without any support. The produced meta-points could serve as meaningful potential keypoints for CAPE. Due to the inevitable gap between inherency and annotation, we finally utilize the identities and details offered by support keypoints to assign and refine meta-points to desired keypoints in query image. In addition, we propose a progressive deformable point decoder and a slacked regression loss for better prediction and supervision. Our novel framework not only reveals the inherency of keypoints but also outperforms existing methods of CAPE. Comprehensive experiments and in-depth studies on large-scale MP-100 dataset demonstrate the effectiveness of our framework.
翻译:类别无关姿态估计(CAPE)旨在通过少量带关键点标注的支持图像,预测任意类别的关键点。现有方法仅依赖支持关键点处提取的特征来预测或精炼查询图像上的关键点,但少量支持特征向量具有局部性,不足以应对CAPE任务。鉴于人类能快速感知任意物体的潜在关键点,我们提出一种基于此类潜在关键点(称为元点)的新型CAPE框架。具体而言,我们维护可学习的嵌入向量以捕获各类关键点的固有信息,这些嵌入能与图像特征图交互,在无需任何支持样本的情况下生成元点。生成的元点可作为CAPE中有意义的潜在关键点。由于固有信息与标注之间存在不可避免的差距,我们最终利用支持关键点提供的身份和细节信息,将元点分配并精炼为查询图像中的目标关键点。此外,我们提出渐进式可变形点解码器和松弛回归损失,以实现更优的预测与监督。该新框架不仅揭示了关键点的固有特性,而且性能超越现有CAPE方法。在大规模MP-100数据集上的全面实验和深度研究验证了我们框架的有效性。