Interaction intention anticipation aims to jointly predict future hand trajectories and interaction hotspots. Existing research often treated trajectory forecasting and interaction hotspots prediction as separate tasks or solely considered the impact of trajectories on interaction hotspots, which led to the accumulation of prediction errors over time. However, a deeper inherent connection exists between hand trajectories and interaction hotspots, which allows for continuous mutual correction between them. Building upon this relationship, a novel Bidirectional prOgressive Transformer (BOT), which introduces a Bidirectional Progressive mechanism into the anticipation of interaction intention is established. Initially, BOT maximizes the utilization of spatial information from the last observation frame through the Spatial-Temporal Reconstruction Module, mitigating conflicts arising from changes of view in first-person videos. Subsequently, based on two independent prediction branches, a Bidirectional Progressive Enhancement Module is introduced to mutually improve the prediction of hand trajectories and interaction hotspots over time to minimize error accumulation. Finally, acknowledging the intrinsic randomness in human natural behavior, we employ a Trajectory Stochastic Unit and a C-VAE to introduce appropriate uncertainty to trajectories and interaction hotspots, respectively. Our method achieves state-of-the-art results on three benchmark datasets Epic-Kitchens-100, EGO4D, and EGTEA Gaze+, demonstrating superior in complex scenarios.
翻译:交互意图预测旨在联合预测未来的手部轨迹与交互热点。现有研究通常将轨迹预测与交互热点预测视为独立任务,或仅考虑轨迹对交互热点的影响,导致预测误差随时间累积。然而,手部轨迹与交互热点之间存在更深层的内在联系,使得两者能够持续相互修正。基于这一关系,本文提出了一种新颖的双向渐进式Transformer(Bidirectional Progressive Transformer, BOT),将双向渐进机制引入交互意图预测中。首先,BOT通过时空重建模块最大化利用最后观测帧的空间信息,缓解第一人称视频中视角变化引发的冲突。随后,基于两条独立预测分支,引入双向渐进增强模块,使手部轨迹与交互热点的预测随时间相互优化,以最小化误差累积。最后,考虑到人类自然行为的内在随机性,我们分别采用轨迹随机单元和条件变分自编码器(C-VAE)为轨迹和交互热点引入适当的随机性。本方法在Epic-Kitchens-100、EGO4D和EGTEA Gaze+三个基准数据集上取得了最优性能,尤其在复杂场景中表现卓越。