We study the problem of human gaze modeling, which aims to generate the gaze patterns a viewer produces while observing a visual stimulus. Gaze is primarily captured through two modalities: continuous eye-tracking trajectories, which describe fine-grained motion dynamics, and discrete scanpaths, which describe high-level fixation structure. Because gaze varies substantially across viewers and trials, we treat this variability as a defining property rather than noise and model gaze as a stochastic generative process. Existing generative gaze models supervise on only one of these two representations in isolation. We hypothesize that trajectories and scanpaths describe gaze at complementary scales and are jointly informative during training, and test this hypothesis through ST-DiffEye, a joint trajectory-scanpath diffusion framework that couples both modalities by concatenating them as an additional raw input channel, requiring no architectural overhead beyond an input and output channel expansion. We further introduce a principled evaluation framework based on the Continuous Ranked Probability Score (CRPS), which generalizes any existing sequence similarity metric into a proper scoring rule that jointly assesses the accuracy and diversity of generated gaze. Experiments on task-driven visual search, covering both target-present and target-absent scenarios, and on free-viewing benchmarks demonstrate state-of-the-art performance. These results, along with detailed ablations, confirm the benefit of joint modeling and the value of distribution-aware evaluation in capturing the intrinsic variability of human gaze. Project webpage: https://st-diffeye.github.io/


翻译:我们研究人类注视建模问题,目标在于生成观察者在观看视觉刺激时产生的注视模式。注视主要通过两种模态捕获:描述细粒度运动动力学的连续眼动追踪轨迹,以及描述高层次注视结构的离散扫视路径。由于注视在不同观察者和试次间存在显著差异,我们将其视为定义性特征而非噪声,并将注视建模为随机生成过程。现有生成式注视模型仅对上述两种表征之一进行监督训练。我们假设轨迹与扫视路径在互补尺度上描述注视,且训练时两者共同携带信息,并通过ST-DiffEye验证这一假设——该联合轨迹-扫视路径扩散框架通过将两种模态拼接为额外原始输入通道实现耦合,除输入输出通道扩展外无需额外架构开销。我们进一步提出基于连续排名概率得分(CRPS)的原理性评估框架,该框架将任意现有序列相似性度量泛化为可同时评估生成注视准确性与多样性的恰当评分规则。涵盖目标存在与目标缺失场景的任务驱动视觉搜索实验,以及自由观看基准测试均展现出最先进性能。这些结果与详细消融研究共同验证了联合建模的优势,以及分布感知评估在捕捉人类注视内在变异性方面的价值。项目网页:https://st-diffeye.github.io/

0
下载
关闭预览

相关内容

扩散模型中的注意力机制:综述
专知会员服务
25+阅读 · 2025年4月10日
标注受限场景下的视觉表征与理解
专知会员服务
14+阅读 · 2025年2月6日
《基于扩散模型的条件图像生成》综述
专知会员服务
44+阅读 · 2024年10月1日
LinkedIn最新《注意力模型》综述论文大全,20页pdf
专知会员服务
138+阅读 · 2020年12月20日
计算机视觉方向简介 | 多目标跟踪算法(附源码)
计算机视觉life
15+阅读 · 2019年6月26日
干货 | 视频显著性目标检测(文末附有完整源码)
计算机视觉战队
14+阅读 · 2019年4月29日
深度学习中的注意力机制
CSDN大数据
24+阅读 · 2017年11月2日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
《多域冲突比较支持模型》60页
专知会员服务
4+阅读 · 8月7日
面向2027年及未来的海军情报改革
专知会员服务
4+阅读 · 8月5日
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员