Domain adaptive pose estimation aims to enable deep models trained on source domain (synthesized) datasets produce similar results on the target domain (real-world) datasets. The existing methods have made significant progress by conducting image-level or feature-level alignment. However, only aligning at a single level is not sufficient to fully bridge the domain gap and achieve excellent domain adaptive results. In this paper, we propose a multi-level domain adaptation aproach, which aligns different domains at the image, feature, and pose levels. Specifically, we first utilize image style transer to ensure that images from the source and target domains have a similar distribution. Subsequently, at the feature level, we employ adversarial training to make the features from the source and target domains preserve domain-invariant characeristics as much as possible. Finally, at the pose level, a self-supervised approach is utilized to enable the model to learn diverse knowledge, implicitly addressing the domain gap. Experimental results demonstrate that significant imrovement can be achieved by the proposed multi-level alignment method in pose estimation, which outperforms previous state-of-the-art in human pose by up to 2.4% and animal pose estimation by up to 3.1% for dogs and 1.4% for sheep.
翻译:域自适应姿态估计旨在使在源域(合成)数据集上训练的深度模型能够对目标域(真实世界)数据集产生相似的结果。现有方法通过图像级或特征级对齐取得了显著进展。然而,仅在单一层次上进行对齐不足以完全弥合域间差距并实现优异的域自适应效果。本文提出了一种多层次域自适应方法,该方法在图像、特征和姿态三个层次上对齐不同域。具体而言,我们首先利用图像风格迁移确保源域和目标域图像具有相似的分布。随后,在特征层面,我们采用对抗训练使得源域和目标域的特征尽可能保持域不变特性。最后,在姿态层面,利用自监督方法使模型学习多样化知识,隐式地解决域差距问题。实验结果表明,所提出的多层次对齐方法在姿态估计方面能带来显著提升,在人脸姿态估计上较先前最先进方法提升高达2.4%,在动物姿态估计中针对狗和羊分别提升高达3.1%和1.4%。