Recognizing facial activity is a well-understood (but non-trivial) computer vision problem. However, reliable solutions require a camera with a good view of the face, which is often unavailable in wearable settings. Furthermore, in wearable applications, where systems accompany users throughout their daily activities, a permanently running camera can be problematic for privacy (and legal) reasons. This work presents an alternative solution based on the fusion of wearable inertial sensors, planar pressure sensors, and acoustic mechanomyography (muscle sounds). The sensors were placed unobtrusively in a sports cap to monitor facial muscle activities related to facial expressions. We present our integrated wearable sensor system, describe data fusion and analysis methods, and evaluate the system in an experiment with thirteen subjects from different cultural backgrounds (eight countries) and both sexes (six women and seven men). In a one-model-per-user scheme and using a late fusion approach, the system yielded an average F1 score of 85.00% for the case where all sensing modalities are combined. With a cross-user validation and a one-model-for-all-user scheme, an F1 score of 79.00% was obtained for thirteen participants (six females and seven males). Moreover, in a hybrid fusion (cross-user) approach and six classes, an average F1 score of 82.00% was obtained for eight users. The results are competitive with state-of-the-art non-camera-based solutions for a cross-user study. In addition, our unique set of participants demonstrates the inclusiveness and generalizability of the approach.
翻译:识别面部活动是一个被充分研究(但并非微不足道)的计算机视觉问题。然而,可靠的解决方案需要摄像头能够清晰捕捉面部,这在可穿戴设备场景中往往无法实现。此外,在可穿戴应用中,当系统伴随用户进行日常活动时,持续运行的摄像头可能因隐私(和法律)原因引发问题。本文提出了一种替代方案,基于可穿戴惯性传感器、平面压力传感器和声学机械肌电图(肌肉声音)的融合。这些传感器被无干扰地嵌入一顶运动帽中,用于监测与面部表情相关的面部肌肉活动。我们介绍了集成的可穿戴传感器系统,描述了数据融合与分析方法,并通过一项涉及来自不同文化背景(八个国家)的十三名受试者(包括六名女性和七名男性)的实验对系统进行了评估。在每位用户单独建模的方案中,采用后期融合方法,当所有传感模态结合时,系统获得的平均F1分数为85.00%。在跨用户验证和所有用户共用单一模型的方案下,十三名参与者(六名女性和七名男性)的F1分数为79.00%。此外,在混合融合(跨用户)方法及六类分类场景中,八名用户的平均F1分数达到82.00%。这些结果与当前最先进的非摄像头解决方案相比具有竞争力,可用于跨用户研究。同时,我们独特的参与者群体证明了该方法的包容性和泛化能力。