In this paper, we introduce EyeEcho, a minimally-obtrusive acoustic sensing system designed to enable glasses to continuously monitor facial expressions. It utilizes two pairs of speakers and microphones mounted on glasses, to emit encoded inaudible acoustic signals directed towards the face, capturing subtle skin deformations associated with facial expressions. The reflected signals are processed through a customized machine-learning pipeline to estimate full facial movements. EyeEcho samples at 83.3 Hz with a relatively low power consumption of 167 mW. Our user study involving 12 participants demonstrates that, with just four minutes of training data, EyeEcho achieves highly accurate tracking performance across different real-world scenarios, including sitting, walking, and after remounting the devices. Additionally, a semi-in-the-wild study involving 10 participants further validates EyeEcho's performance in naturalistic scenarios while participants engage in various daily activities. Finally, we showcase EyeEcho's potential to be deployed on a commercial-off-the-shelf (COTS) smartphone, offering real-time facial expression tracking.
翻译:本文介绍EyeEcho,一种最小侵入性的声学感知系统,旨在使眼镜能够连续监测面部表情。该系统利用安装在眼镜上的两对扬声器和麦克风,向面部发射编码后的不可听见声学信号,捕捉与面部表情相关的细微皮肤形变。反射信号通过定制化机器学习流水线进行处理,以估计完整的面部运动。EyeEcho的采样率为83.3 Hz,功耗较低(167 mW)。一项涉及12名参与者的用户研究表明,仅需四分钟训练数据,EyeEcho即可在不同真实场景(包括静坐、行走及设备重新佩戴后)中实现高精度追踪性能。此外,一项涉及10名参与者的半野生环境研究进一步验证了EyeEcho在参与者进行多种日常活动时的自然场景表现。最后,我们展示了EyeEcho在商用现成(COTS)智能手机上部署的潜力,可提供实时面部表情追踪功能。