Self-supervised landmark estimation is a challenging task that demands the formation of locally distinct feature representations to identify sparse facial landmarks in the absence of annotated data. To tackle this task, existing state-of-the-art (SOTA) methods (1) extract coarse features from backbones that are trained with instance-level self-supervised learning (SSL) paradigms, which neglect the dense prediction nature of the task, (2) aggregate them into memory-intensive hypercolumn formations, and (3) supervise lightweight projector networks to naively establish full local correspondences among all pairs of spatial features. In this paper, we introduce SCE-MAE, a framework that (1) leverages the MAE, a region-level SSL method that naturally better suits the landmark prediction task, (2) operates on the vanilla feature map instead of on expensive hypercolumns, and (3) employs a Correspondence Approximation and Refinement Block (CARB) that utilizes a simple density peak clustering algorithm and our proposed Locality-Constrained Repellence Loss to directly hone only select local correspondences. We demonstrate through extensive experiments that SCE-MAE is highly effective and robust, outperforming existing SOTA methods by large margins of approximately 20%-44% on the landmark matching and approximately 9%-15% on the landmark detection tasks.
翻译:自监督关键点估计是一项具有挑战性的任务,它需要在缺乏标注数据的情况下,形成局部独特的特征表示以识别稀疏的面部关键点。为应对此任务,现有的最先进方法(1)从通过实例级自监督学习范式训练的主干网络中提取粗糙特征,这忽略了任务的密集预测特性;(2)将这些特征聚合为内存密集型的超列形式;以及(3)监督轻量级投影器网络以朴素地建立所有空间特征对之间的完整局部对应关系。在本文中,我们提出了SCE-MAE框架,该框架(1)利用了MAE这一区域级自监督学习方法,其天然更适合关键点预测任务;(2)在原始特征图上操作,而非昂贵的超列上;以及(3)采用了一个对应关系近似与细化模块,该模块利用简单的密度峰值聚类算法和我们提出的局部性约束排斥损失,来直接仅针对选定的局部对应关系进行优化。我们通过大量实验证明,SCE-MAE非常有效且鲁棒,在关键点匹配任务上以约20%-44%的显著优势超越现有最先进方法,在关键点检测任务上以约9%-15%的优势超越现有方法。