Understanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit temporal information for robust estimation. Noticing the different temporal granularity of and the semantic correlation between hand pose estimation and action recognition, we build a network hierarchy with two cascaded transformer encoders, where the first one exploits the short-term temporal cue for hand pose estimation, and the latter aggregates per-frame pose and object information over a longer time span to recognize the action. Our approach achieves competitive results on two first-person hand action benchmarks, namely FPHA and H2O. Extensive ablation studies verify our design choices.
翻译:理解自我中心RGB视频中的动态手部动作与行为是一项基础但具有挑战性的任务,主要源于自遮挡和歧义性。为解决遮挡与歧义问题,我们提出了一种基于变换器的框架,利用时序信息实现鲁棒估计。注意到手部姿态估计与动作识别在时间粒度上的差异及语义相关性,我们构建了包含两个级联变换器编码器的网络层级结构:第一个编码器利用短期时序线索进行手部姿态估计;第二个编码器则在更长时间跨度内聚合每帧的姿态及物体信息以识别动作。我们的方法在FPHA和H2O两个第一人称手部动作基准数据集上取得了具有竞争力的结果。广泛的消融实验验证了我们的设计选择。