In this letter, we propose a new method, Multi-Clue Gaze (MCGaze), to facilitate video gaze estimation via capturing spatial-temporal interaction context among head, face, and eye in an end-to-end learning way, which has not been well concerned yet. The main advantage of MCGaze is that the tasks of clue localization of head, face, and eye can be solved jointly for gaze estimation in a one-step way, with joint optimization to seek optimal performance. During this, spatial-temporal context exchange happens among the clues on the head, face, and eye. Accordingly, the final gazes obtained by fusing features from various queries can be aware of global clues from heads and faces, and local clues from eyes simultaneously, which essentially leverages performance. Meanwhile, the one-step running way also ensures high running efficiency. Experiments on the challenging Gaze360 dataset verify the superiority of our proposition. The source code will be released at https://github.com/zgchen33/MCGaze.
翻译:在本文中,我们提出了一种新方法——多线索视线估计(MCGaze),通过以端到端学习方式捕捉头部、面部和眼部之间的时空交互上下文来促进视频视线估计,这一问题此前尚未得到充分关注。MCGaze的主要优势在于:头部、面部和眼部的线索定位任务可以联合求解,以一步到位的方式实现视线估计,并通过联合优化寻求最佳性能。在此过程中,时空上下文在头部、面部和眼部的线索之间进行交换。因此,通过融合来自不同查询的特征所获得的最终视线能够同时感知来自头部和面部的全局线索以及来自眼部的局部线索,这从根本上提升了性能。同时,一步式运行方式也确保了较高的运行效率。在具有挑战性的Gaze360数据集上的实验验证了本文所提方法的优越性。源代码将在https://github.com/zgchen33/MCGaze发布。