While zero-shot appearance-based 3D gaze estimation offers significant cost-efficiency by directly mapping RGB images to gaze vectors, its reliability in Human-Robot Interaction (HRI) settings remains uncertain. Existing benchmarks frequently overlook fundamental HRI conditions, such as dynamic camera viewpoints and moving targets in video. Furthermore, current cross-dataset evaluations often suffer from a complexity gap, where methods trained on diverse datasets are tested on significantly smaller and less varied sets, failing to assess true robustness. To bridge these gaps, we introduce Gaze4HRI, a large-scale dataset (50+ subjects, 3,000+ videos, 600,000+ frames) designed to evaluate state-of-the-art performance against critical HRI variables: illumination, head-gaze conflict, as well as the motion of camera and gaze target in video. Our benchmark reveals that all evaluated methods fail in at least one condition, identifying steeply-downward gaze as a universal failure point. Notably, PureGaze trained on the ETH-X-Gaze dataset uniquely maintains resilience across all other conditions. These results challenge the recent focus in the literature on complex spatial-temporal modeling and Transformer-based architectures. Instead, our findings suggest that extensive data diversity, as exemplified by the ETH-X-Gaze dataset, serves as the primary driver of zero-shot robustness in unconstrained environments, while resilience-enhancing frameworks, such as PureGaze's self-adversarial loss for gaze feature purification, provide a substantial further improvement. Ultimately, this study establishes a rigorous benchmark that provides practical guidelines for practitioners as well as reshaping future research. The dataset and codes are available at https://gazeforhri.github.io.


翻译:尽管基于外观的零样本三维凝视估计通过直接将RGB图像映射为凝视向量具有显著的成本效益,但其在人机交互场景中的可靠性仍不明确。现有基准测试常忽视人机交互的基本条件,例如视频中的动态相机视角与移动目标。此外,跨数据集评估常因复杂度差异而失效——基于多样化数据集训练的方法被测试于规模更小、变化更少的数据集,难以评估真实鲁棒性。为弥合这些差距,我们提出Gaze4HRI——一个大规模数据集(含50余名被试、3000余段视频、60万余帧),旨在针对人机交互关键变量(光照、头姿-凝视冲突、以及视频中相机与凝视目标的运动)评估最先进方法的性能。基准测试表明:所有评估方法至少在某项条件中失效,其中陡峭向下凝视被识别为通用失效点。值得注意的是,基于ETH-X-Gaze数据集训练的PureGaze在其余所有条件下均保持鲁棒性。这些结果挑战了近期文献对复杂时空建模与Transformer架构的重视。相反,我们的发现表明:以ETH-X-Gaze数据集为代表的广泛数据多样性,是开放环境中零样本鲁棒性的主要驱动因素;而诸如PureGaze用于凝视特征纯化的自对抗损失等韧性增强框架,则可提供进一步的显著改进。最终,本研究建立了严格的基准测试,既为从业者提供实用指南,亦为未来研究方向重塑奠定基础。数据集与代码发布于https://gazeforhri.github.io。

0
下载
关闭预览

相关内容

数据集,又称为资料集、数据集合或资料集合,是一种由数据所组成的集合。
Data set(或dataset)是一个数据的集合,通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量,如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数,该数据集的数据可能包括一个或多个成员。
迈向深度基础模型:基于视觉的深度估计最新趋势
专知会员服务
24+阅读 · 2025年7月16日
《AI生成视频评估综述》
专知会员服务
28+阅读 · 2024年10月30日
基于深度学习的物体姿态估计综述
专知会员服务
27+阅读 · 2024年5月15日
【CVPR2022】GaTector:凝视对象预测的统一框架
专知会员服务
10+阅读 · 2022年3月24日
【AAAI2022】基于特征纯化的视线估计算法
专知会员服务
10+阅读 · 2022年2月11日
最新《深度学习人体姿态估计》综述论文,26页pdf
专知会员服务
41+阅读 · 2020年12月29日
深度学习人体姿态估计算法综述
AI前线
25+阅读 · 2019年5月19日
报名 | 让机器读懂你的意图——人体姿态估计入门
人工智能头条
10+阅读 · 2017年9月19日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
28+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
从采集到决策:美军视角下的战术情报范式重构
专知会员服务
0+阅读 · 19分钟前
《履带式无人地面战车技术发展现状》
专知会员服务
1+阅读 · 今天1:46
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
相关基金
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
28+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员