With recent advancements in computer vision as well as machine learning (ML), video-based at-home exercise evaluation systems have become a popular topic of current research. However, performance depends heavily on the amount of available training data. Since labeled datasets specific to exercising are rare, we propose a method that makes use of the abundance of fitness videos available online. Specifically, we utilize the advantage that videos often not only show the exercises, but also provide language as an additional source of information. With push-ups as an example, we show that through the analysis of subtitle data using natural language processing (NLP), it is possible to create a labeled (irrelevant, relevant correct, relevant incorrect) dataset containing relevant information for pose analysis. In particular, we show that irrelevant clips ($n=332$) have significantly different joint visibility values compared to relevant clips ($n=298$). Inspecting cluster centroids also show different poses for the different classes.
翻译:随着计算机视觉与机器学习的最新进展,基于视频的家庭运动评估系统已成为当前研究热点。然而,系统性能高度依赖训练数据的规模。由于专门针对运动训练的标注数据集稀缺,本文提出了一种利用互联网上海量健身视频的方法。具体而言,我们利用了视频不仅展示运动过程、还通过语言提供附加信息这一优势。以俯卧撑为例,我们证明通过自然语言处理(NLP)分析字幕数据,能够生成包含姿态分析相关信息的标注数据集(分为无关、相关正确、相关错误三类)。研究表明,无关片段(n=332)与相关片段(n=298)的关节可见性值存在显著差异。对聚类中心点的分析也显示不同类别的姿态模式存在差异。