Research into the detection of human activities from wearable sensors is a highly active field, benefiting numerous applications, from ambulatory monitoring of healthcare patients via fitness coaching to streamlining manual work processes. We present an empirical study that compares 4 different commonly used annotation methods utilized in user studies that focus on in-the-wild data. These methods can be grouped in user-driven, in situ annotations - which are performed before or during the activity is recorded - and recall methods - where participants annotate their data in hindsight at the end of the day. Our study illustrates that different labeling methodologies directly impact the annotations' quality, as well as the capabilities of a deep learning classifier trained with the data respectively. We noticed that in situ methods produce less but more precise labels than recall methods. Furthermore, we combined an activity diary with a visualization tool that enables the participant to inspect and label their activity data. Due to the introduction of such a tool were able to decrease missing annotations and increase the annotation consistency, and therefore the F1-score of the deep learning model by up to 8% (ranging between 82.1 and 90.4% F1-score). Furthermore, we discuss the advantages and disadvantages of the methods compared in our study, the biases they may could introduce and the consequences of their usage on human activity recognition studies and as well as possible solutions.
翻译:基于可穿戴传感器的人类活动检测研究是一个高度活跃的领域,惠及从医疗患者的动态监测、健身指导到简化手工工作流程等众多应用。本文通过实证研究,比较了针对野外数据的用户研究中四种常用标注方法。这些方法可分为用户驱动的实况标注(在活动记录前或记录时执行)和回忆型标注(参与者在当天结束时回顾性标注数据)。研究表明,不同标注方法不仅直接影响标注质量,还会影响基于该数据训练的深度学习分类器性能。我们注意到,实况方法产生的标注数量较少但精度更高。此外,我们将活动日记与可视化工具结合,使参与者能够查看并标注其活动数据。该工具的引入减少了缺失标注并提高了标注一致性,从而使深度学习模型的F1分数提升最高达8%(F1分数范围从82.1%至90.4%)。进一步地,本文讨论了研究中比较方法的优缺点、可能引入的偏差、对人类活动识别研究的影响以及可能的解决方案。