With the escalated demand of human-machine interfaces for intelligent systems, development of gaze controlled system have become a necessity. Gaze, being the non-intrusive form of human interaction, is one of the best suited approach. Appearance based deep learning models are the most widely used for gaze estimation. But the performance of these models is entirely influenced by the size of labeled gaze dataset and in effect affects generalization in performance. This paper aims to develop a semi-supervised contrastive learning framework for estimation of gaze direction. With a small labeled gaze dataset, the framework is able to find a generalized solution even for unseen face images. In this paper, we have proposed a new contrastive loss paradigm that maximizes the similarity agreement between similar images and at the same time reduces the redundancy in embedding representations. Our contrastive regression framework shows good performance in comparison to several state of the art contrastive learning techniques used for gaze estimation.
翻译:随着智能系统对人机交互界面的需求日益增长,开发基于眼动控制的系统已成为必然趋势。眼动作为一种非侵入式的人机交互方式,是最为适用的方法之一。基于外观的深度学习模型是眼动估计中最广泛使用的方法,但这些模型的性能完全受标注眼动数据集规模的影响,进而影响其泛化性能。本文旨在开发一种用于眼动方向估计的半监督对比学习框架。即使仅使用少量标注眼动数据集,该框架也能为未见人脸图像找到泛化解。本文提出了一种新的对比损失范式,该范式在最大化相似图像间一致性的同时,减少了嵌入表示中的冗余。与当前用于眼动估计的多种先进对比学习技术相比,我们的对比回归框架展现出良好的性能。