Most existing attention prediction research focuses on salient instances like humans and objects. However, the more complex interaction-oriented attention, arising from the comprehension of interactions between instances by human observers, remains largely unexplored. This is equally crucial for advancing human-machine interaction and human-centered artificial intelligence. To bridge this gap, we first collect a novel gaze fixation dataset named IG, comprising 530,000 fixation points across 740 diverse interaction categories, capturing visual attention during human observers cognitive processes of interactions. Subsequently, we introduce the zero-shot interaction-oriented attention prediction task ZeroIA, which challenges models to predict visual cues for interactions not encountered during training. Thirdly, we present the Interactive Attention model IA, designed to emulate human observers cognitive processes to tackle the ZeroIA problem. Extensive experiments demonstrate that the proposed IA outperforms other state-of-the-art approaches in both ZeroIA and fully supervised settings. Lastly, we endeavor to apply interaction-oriented attention to the interaction recognition task itself. Further experimental results demonstrate the promising potential to enhance the performance and interpretability of existing state-of-the-art HOI models by incorporating real human attention data from IG and attention labels generated by IA.
翻译:现有注意力预测研究大多聚焦于人和物等显著实例。然而,由人类观察者对实例间交互理解产生的更复杂的交互导向型注意力仍鲜有探索。这对于推进人机交互和以人为中心的人工智能同样至关重要。为填补这一空白,我们首先收集了一个名为IG的新型注视数据集,包含740种不同交互类别中的53万个注视点,捕捉了人类观察者在交互认知过程中的视觉注意力。随后,我们提出了零样本交互导向型注意力预测任务ZeroIA,该任务要求模型为训练中未见过的交互预测视觉线索。第三,我们提出了交互注意力模型IA,该模型旨在模拟人类观察者的认知过程以解决ZeroIA问题。大量实验表明,所提出的IA在ZeroIA和全监督设置下均优于其他最先进方法。最后,我们尝试将交互导向型注意力应用于交互识别任务本身。进一步的实验结果展示了通过将IG的真实人类注意力数据和IA生成的注意力标签融入现有最先进HOI模型,有望增强其性能与可解释性。