Performance analyses based on videos are commonly used by coaches of athletes in various sports disciplines. In individual sports, these analyses mainly comprise the body posture. This paper focuses on the disciplines of triple, high, and long jump, which require fine-grained locations of the athlete's body. Typical human pose estimation datasets provide only a very limited set of keypoints, which is not sufficient in this case. Therefore, we propose a method to detect arbitrary keypoints on the whole body of the athlete by leveraging the limited set of annotated keypoints and auto-generated segmentation masks of body parts. Evaluations show that our model is capable of detecting keypoints on the head, torso, hands, feet, arms, and legs, including also bent elbows and knees. We analyze and compare different techniques to encode desired keypoints as the model's input and their embedding for the Transformer backbone.
翻译:基于视频的表现分析广泛应用于各运动项目教练员的训练指导中。在个人运动中,此类分析主要涉及运动员的身体姿态。本文聚焦于三级跳远、跳高及跳远项目,这些项目需要对运动员身体部位进行精细定位。现有典型人体姿态估计数据集仅提供非常有限的关键点集合,难以满足此类需求。为此,我们提出一种方法,通过利用有限标注关键点与自动生成的躯干分割掩码,检测运动员全身任意关键点。评估结果表明,该模型能够检测头部、躯干、手部、足部、手臂及腿部关键点,包括弯曲的肘关节与膝关节。我们分析并比较了将目标关键点编码为模型输入的不同技术及其在Transformer主干网络中的嵌入方式。