VR Facial Animation is necessary in applications requiring clear view of the face, even though a VR headset is worn. In our case, we aim to animate the face of an operator who is controlling our robotic avatar system. We propose a real-time capable pipeline with very fast adaptation for specific operators. In a quick enrollment step, we capture a sequence of source images from the operator without the VR headset which contain all the important operator-specific appearance information. During inference, we then use the operator keypoint information extracted from a mouth camera and two eye cameras to estimate the target expression and head pose, to which we map the appearance of a source still image. In order to enhance the mouth expression accuracy, we dynamically select an auxiliary expression frame from the captured sequence. This selection is done by learning to transform the current mouth keypoints into the source camera space, where the alignment can be determined accurately. We, furthermore, demonstrate an eye tracking pipeline that can be trained in less than a minute, a time efficient way to train the whole pipeline given a dataset that includes only complete faces, show exemplary results generated by our method, and discuss performance at the ANA Avatar XPRIZE semifinals.
翻译:VR面部动画在需要清晰观察面部表情的应用中不可或缺,即使用户佩戴着VR头显。本研究旨在为操控机器人化身系统的人类操作员生成面部动画。我们提出一种支持实时运行且可快速适配特定操作员的流程。在快速注册阶段,我们采集操作员未佩戴VR头显时的源图像序列,该序列包含所有重要的操作员专属外观信息。推理阶段,我们利用从嘴部摄像头和两个眼部摄像头提取的操作员关键点信息,估计目标表情与头部姿态,从而将源静态图像的外观映射至目标状态。为提升嘴部表情精度,我们动态从已采集序列中选取辅助表情帧:通过学习将当前嘴部关键点变换至源摄像机空间,实现精准对齐。此外,我们提出可在1分钟内完成训练的眼球追踪流程,并展示了一种高效训练完整流程的方法(基于仅含完整面部的数据集)。最后,我们呈现方法生成的示例结果,并讨论ANA Avatar XPRIZE半决赛中的性能表现。