Gaze-annotated facial data is crucial for training deep neural networks (DNNs) for gaze estimation. However, obtaining these data is labor-intensive and requires specialized equipment due to the challenge of accurately annotating the gaze direction of a subject. In this work, we present a generative framework to create annotated gaze data by leveraging the benefits of labeled and unlabeled data sources. We propose a Gaze-aware Compositional GAN that learns to generate annotated facial images from a limited labeled dataset. Then we transfer this model to an unlabeled data domain to take advantage of the diversity it provides. Experiments demonstrate our approach's effectiveness in generating within-domain image augmentations in the ETH-XGaze dataset and cross-domain augmentations in the CelebAMask-HQ dataset domain for gaze estimation DNN training. We also show additional applications of our work, which include facial image editing and gaze redirection.
翻译:注视标注的面部数据对于训练深度神经网络(DNN)进行视线估计至关重要。然而,由于准确标注主体注视方向具有挑战性,获取此类数据不仅劳动密集,还需要专用设备。在本工作中,我们提出一个生成框架,通过利用标注与未标注数据源的优势,来创建带标注的注视数据。我们提出了一种注视感知的组合生成对抗网络,该网络学习从有限的标注数据集中生成带标注的面部图像。随后,我们将此模型迁移到一个未标注的数据域,以利用该域提供的多样性。实验证明,我们的方法在ETH-XGaze数据集中生成域内图像增强,以及在CelebAMask-HQ数据集域中生成跨域增强以用于视线估计DNN训练方面具有有效性。我们还展示了本工作的其他应用,包括面部图像编辑与视线重定向。