Event camera shows great potential in 3D hand pose estimation, especially addressing the challenges of fast motion and high dynamic range in a low-power way. However, due to the asynchronous differential imaging mechanism, it is challenging to design event representation to encode hand motion information especially when the hands are not moving (causing motion ambiguity), and it is infeasible to fully annotate the temporally dense event stream. In this paper, we propose EvHandPose with novel hand flow representations in Event-to-Pose module for accurate hand pose estimation and alleviating the motion ambiguity issue. To solve the problem under sparse annotation, we design contrast maximization and edge constraints in Pose-to-IWE (Image with Warped Events) module and formulate EvHandPose in a self-supervision framework. We further build EvRealHands, the first large-scale real-world event-based hand pose dataset on several challenging scenes to bridge the domain gap due to relying on synthetic data and facilitate future research. Experiments on EvRealHands demonstrate that EvHandPose outperforms previous event-based method under all evaluation scenes with 15 $\sim$ 20 mm lower MPJPE and achieves accurate and stable hand pose estimation in fast motion and strong light scenes compared with RGB-based methods. Furthermore, EvHandPose demonstrates 3D hand pose estimation at 120 fps or higher.
翻译:事件相机在3D手部姿态估计中展现出巨大潜力,尤其能以低功耗方式应对快速运动和高动态范围等挑战。然而,由于其异步差分成像机制,设计能编码手部运动信息的事件表征极具挑战性——特别是在手部静止时(会导致运动歧义),且对时间密度极高的事件流进行完整标注不可行。本文提出EvHandPose方法,通过事件-姿态模块中的新型手部流表征实现精确手部姿态估计,并缓解运动歧义问题。为解决稀疏标注下的难题,我们在姿态-加权事件图像模块中设计了对比度最大化和边缘约束,并构建了自监督框架下的EvHandPose。进一步地,我们建立了EvRealHands——首个覆盖多个挑战性场景的大规模真实世界事件驱动手部姿态数据集,以弥合依赖合成数据带来的领域差异,并推动未来研究。在EvRealHands上的实验表明,EvHandPose在所有评估场景中均优于先前基于事件的方法,MPJPE降低15~20毫米;与基于RGB的方法相比,能在快速运动和强光场景中实现精确稳定的手部姿态估计。此外,EvHandPose可实现120帧/秒以上的3D手部姿态估计。