Understanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit temporal information for robust estimation. Noticing the different temporal granularity of and the semantic correlation between hand pose estimation and action recognition, we build a network hierarchy with two cascaded transformer encoders, where the first one exploits the short-term temporal cue for hand pose estimation, and the latter aggregates per-frame pose and object information over a longer time span to recognize the action. Our approach achieves competitive results on two first-person hand action benchmarks, namely FPHA and H2O. Extensive ablation studies verify our design choices.
翻译:从头部视角RGB视频中理解动态手部动作与行为是一项基础但充满挑战的任务,主要受自遮挡与模糊性影响。为应对遮挡与模糊问题,我们提出基于Transformer的框架,通过利用时序信息实现鲁棒估计。注意到手部姿态估计与动作识别在时序粒度上的差异及语义关联,我们构建了一个包含两个级联Transformer编码器的网络层级结构:前者利用短时序线索进行手部姿态估计,后者在更长时间跨度内聚合每帧的姿态与物体信息以识别动作。该方法在两个第一人称手部动作基准数据集(FPHA与H2O)上取得了具有竞争力的结果。广泛的消融实验验证了我们的设计选择。