We introduce HARPER, a novel dataset for 3D body pose estimation and forecast in dyadic interactions between users and \spot, the quadruped robot manufactured by Boston Dynamics. The key-novelty is the focus on the robot's perspective, i.e., on the data captured by the robot's sensors. These make 3D body pose analysis challenging because being close to the ground captures humans only partially. The scenario underlying HARPER includes 15 actions, of which 10 involve physical contact between the robot and users. The Corpus contains not only the recordings of the built-in stereo cameras of Spot, but also those of a 6-camera OptiTrack system (all recordings are synchronized). This leads to ground-truth skeletal representations with a precision lower than a millimeter. In addition, the Corpus includes reproducible benchmarks on 3D Human Pose Estimation, Human Pose Forecasting, and Collision Prediction, all based on publicly available baseline approaches. This enables future HARPER users to rigorously compare their results with those we provide in this work.
翻译:我们提出了HARPER——一个面向用户与波士顿动力公司四足机器人Spot之间二元交互的三维人体姿态估计与预测的新型数据集。其核心创新在于聚焦机器人视角,即通过机器人自身传感器捕获的数据。这些数据使三维人体姿态分析面临挑战,因为机器人近地视场仅能捕捉人体局部。HARPER数据集包含15种动作场景,其中10种涉及机器人与用户的物理接触。该语料库不仅收录了Spot内置立体摄像头的记录数据,还包含六摄像头OptiTrack系统的同步记录,从而获得精度低于一毫米的真实骨骼表征。此外,该语料库还提供了基于公开基线方法的可复现基准测试,涵盖三维人体姿态估计、人体姿态预测和碰撞预测三大任务。这使未来的HARPER用户能够严格比较其研究结果与本文提供的基准数据。