Robots are becoming increasingly integrated into our lives, assisting us in various tasks. To ensure effective collaboration between humans and robots, it is essential that they understand our intentions and anticipate our actions. In this paper, we propose a Human-Object Interaction (HOI) anticipation framework for collaborative robots. We propose an efficient and robust transformer-based model to detect and anticipate HOIs from videos. This enhanced anticipation empowers robots to proactively assist humans, resulting in more efficient and intuitive collaborations. Our model outperforms state-of-the-art results in HOI detection and anticipation in VidHOI dataset with an increase of 1.76% and 1.04% in mAP respectively while being 15.4 times faster. We showcase the effectiveness of our approach through experimental results in a real robot, demonstrating that the robot's ability to anticipate HOIs is key for better Human-Robot Interaction. More information can be found on our project webpage: https://evm7.github.io/HOI4ABOT_page/
翻译:机器人正日益融入我们的生活,协助我们完成各类任务。为确保人机高效协作,机器人必须理解人类意图并预测人类行为。本文提出一种面向协作机器人的物体交互预测框架。我们设计了一种高效稳健的基于Transformer的模型,用于从视频中检测并预测物体交互行为。这种增强的预测能力使机器人能够主动协助人类,实现更高效、更直观的协作。在VidHOI数据集上,我们的模型在物体交互检测与预测任务中分别以1.76%和1.04%的mAP提升超越现有最优方法,同时推理速度提升15.4倍。通过真实机器人实验验证了方法的有效性,证明机器人预测物体交互的能力是人机交互优化的关键。更多信息可访问项目网页:https://evm7.github.io/HOI4ABOT_page/