Performance analyses based on videos are commonly used by coaches of athletes in various sports disciplines. In individual sports, these analyses mainly comprise the body posture. This paper focuses on the disciplines of triple, high, and long jump, which require fine-grained locations of the athlete's body. Typical human pose estimation datasets provide only a very limited set of keypoints, which is not sufficient in this case. Therefore, we propose a method to detect arbitrary keypoints on the whole body of the athlete by leveraging the limited set of annotated keypoints and auto-generated segmentation masks of body parts. Evaluations show that our model is capable of detecting keypoints on the head, torso, hands, feet, arms, and legs, including also bent elbows and knees. We analyze and compare different techniques to encode desired keypoints as the model's input and their embedding for the Transformer backbone.
翻译:基于视频的竞技表现分析是各运动项目教练普遍采用的方法。在个人项目中,此类分析主要涉及身体姿态。本文聚焦三级跳、跳高和跳远项目,这些项目需要对运动员身体进行精细定位。典型的人体姿态估计数据集仅提供非常有限的关键点集,在此情况下不足以满足需求。因此,我们提出一种方法,通过利用有限标注关键点集和自动生成的身体部位分割掩码,检测运动员全身任意关键点。评估表明,我们的模型能够检测头部、躯干、双手、双足、双臂和双腿上的关键点,包括弯曲的肘部和膝盖。我们分析并比较了将期望关键点编码为模型输入的不同技术及其在Transformer骨干网络中的嵌入方式。