Deep learning has bolstered gaze estimation techniques, but real-world deployment has been impeded by inadequate training datasets. This problem is exacerbated by both hardware-induced variations in eye images and inherent biological differences across the recorded participants, leading to both feature and pixel-level variance that hinders the generalizability of models trained on specific datasets. While synthetic datasets can be a solution, their creation is both time and resource-intensive. To address this problem, we present a framework called Light Eyes or "LEyes" which, unlike conventional photorealistic methods, only models key image features required for video-based eye tracking using simple light distributions. LEyes facilitates easy configuration for training neural networks across diverse gaze-estimation tasks. We demonstrate that models trained using LEyes are consistently on-par or outperform other state-of-the-art algorithms in terms of pupil and CR localization across well-known datasets. In addition, a LEyes trained model outperforms the industry standard eye tracker using significantly more cost-effective hardware. Going forward, we are confident that LEyes will revolutionize synthetic data generation for gaze estimation models, and lead to significant improvements of the next generation video-based eye trackers.
翻译:深度学习推动了注视估计技术的发展,但实际部署因训练数据集不足而受阻。这一问题因硬件导致的眼图像差异及被记录参与者固有的生物学差异而加剧,导致特征层面和像素层面的双重变异,从而削弱了在特定数据集上训练的模型的泛化能力。虽然合成数据集可作为解决方案,但其生成过程既耗时又消耗大量资源。为解决此问题,我们提出名为"Light Eyes"(简称"LEyes")的框架——与传统的逼真渲染方法不同,该框架仅通过简单的光分布建模视频眼动追踪所需的关键图像特征。LEyes支持轻松配置以训练适用于多种注视估计任务的神经网络。实验证明,利用LEyes训练的模型在瞳孔与角膜反射(CR)定位任务中,在多个知名数据集上持续达到或超越当前最优算法。此外,基于LEyes训练的模型在使用成本显著更低的硬件时,仍优于行业标准级眼动仪。展望未来,我们坚信LEyes将彻底变革注视估计模型的合成数据生成方式,并推动下一代视频眼动追踪技术的重大突破。