Head pose estimation has become a crucial area of research in computer vision given its usefulness in a wide range of applications, including robotics, surveillance, or driver attention monitoring. One of the most difficult challenges in this field is managing head occlusions that frequently take place in real-world scenarios. In this paper, we propose a novel and efficient framework that is robust in real world head occlusion scenarios. In particular, we propose an unsupervised latent embedding clustering with regression and classification components for each pose angle. The model optimizes latent feature representations for occluded and non-occluded images through a clustering term while improving fine-grained angle predictions. Experimental evaluation on in-the-wild head pose benchmark datasets reveal competitive performance in comparison to state-of-the-art methodologies with the advantage of having a significant data reduction. We observe a substantial improvement in occluded head pose estimation. Also, an ablation study is conducted to ascertain the impact of the clustering term within our proposed framework.
翻译:人脸姿态估计因其在机器人、监控或驾驶员注意力监测等广泛应用中的重要性,已成为计算机视觉领域的核心研究方向。该领域最具挑战性的难题之一,是真实场景中频繁出现的面部遮挡问题。本文提出一种新颖且高效的框架,可在真实人脸遮挡场景中保持鲁棒性。具体而言,我们针对每个姿态角度提出了一种结合回归与分类分量的无监督隐式嵌入聚类方法。该模型通过聚类项优化遮挡与非遮挡图像的隐式特征表示,同时提升角度预测的细粒度精度。在真实场景人脸姿态基准数据集上的实验评估表明,与现有最优方法相比,本方法在保持竞争力性能的同时具有显著的数据缩减优势。我们观察到遮挡人脸姿态估计性能得到实质性提升。此外,通过消融研究验证了聚类项对所提框架的影响。