3D hand tracking from a monocular video is a very challenging problem due to hand interactions, occlusions, left-right hand ambiguity, and fast motion. Most existing methods rely on RGB inputs, which have severe limitations under low-light conditions and suffer from motion blur. In contrast, event cameras capture local brightness changes instead of full image frames and do not suffer from the described effects. Unfortunately, existing image-based techniques cannot be directly applied to events due to significant differences in the data modalities. In response to these challenges, this paper introduces the first framework for 3D tracking of two fast-moving and interacting hands from a single monocular event camera. Our approach tackles the left-right hand ambiguity with a novel semi-supervised feature-wise attention mechanism and integrates an intersection loss to fix hand collisions. To facilitate advances in this research domain, we release a new synthetic large-scale dataset of two interacting hands, Ev2Hands-S, and a new real benchmark with real event streams and ground-truth 3D annotations, Ev2Hands-R. Our approach outperforms existing methods in terms of the 3D reconstruction accuracy and generalises to real data under severe light conditions.
翻译:单目视频中的三维手部跟踪因手部交互、遮挡、左右手混淆及快速运动而极具挑战性。现有方法多依赖RGB输入,在弱光条件下存在严重局限且易受运动模糊影响。相比之下,事件相机仅捕捉局部亮度变化而非完整图像帧,不受所述效应影响。然而,由于数据模态的显著差异,现有基于图像的技术无法直接应用于事件数据。针对上述挑战,本文首次提出从单目事件相机对两个快速运动且交互的手部进行三维跟踪的框架。本方法通过新颖的半监督特征级注意力机制解决左右手混淆问题,并整合交叉损失以修复手部碰撞。为促进该研究领域的发展,我们发布了两个新数据集:大规模合成交互手数据集Ev2Hands-S,以及包含真实事件流与三维标注真值的真实基准数据集Ev2Hands-R。本方法在三维重建精度上优于现有方法,并能泛化至极端光照条件下的真实数据。