In this work we study the benefits of using tracking and 3D poses for action recognition. To achieve this, we take the Lagrangian view on analysing actions over a trajectory of human motion rather than at a fixed point in space. Taking this stand allows us to use the tracklets of people to predict their actions. In this spirit, first we show the benefits of using 3D pose to infer actions, and study person-person interactions. Subsequently, we propose a Lagrangian Action Recognition model by fusing 3D pose and contextualized appearance over tracklets. To this end, our method achieves state-of-the-art performance on the AVA v2.2 dataset on both pose only settings and on standard benchmark settings. When reasoning about the action using only pose cues, our pose model achieves +10.0 mAP gain over the corresponding state-of-the-art while our fused model has a gain of +2.8 mAP over the best state-of-the-art model. Code and results are available at: https://brjathu.github.io/LART
翻译:本研究探讨了利用跟踪与3D姿态信息进行动作识别的优势。为实现此目标,我们采用拉格朗日视角分析沿着人体运动轨迹的动作,而非固定空间点的信息。这种立场允许我们利用人体的轨迹片段(tracklets)来预测其动作。基于此,我们首先展示了利用3D姿态推断动作的益处,并研究了人与人之间的交互。随后,我们提出了一种拉格朗日动作识别模型,该模型通过融合轨迹片段上的3D姿态与情境化外观特征实现。最终,我们的方法在AVA v2.2数据集上的纯姿态设置与标准基准设置中均达到了最先进性能。在仅依赖姿态线索进行动作推理时,我们的姿态模型相比相应最先进方法获得了+10.0 mAP的提升,而融合模型相比最佳最先进模型获得了+2.8 mAP的增长。代码与结果详见:https://brjathu.github.io/LART