Predicting high-fidelity future human poses, from a historically observed sequence, is decisive for intelligent robots to interact with humans. Deep end-to-end learning approaches, which typically train a generic pre-trained model on external datasets and then directly apply it to all test samples, emerge as the dominant solution to solve this issue. Despite encouraging progress, they remain non-optimal, as the unique properties (e.g., motion style, rhythm) of a specific sequence cannot be adapted. More generally, at test-time, once encountering unseen motion categories (out-of-distribution), the predicted poses tend to be unreliable. Motivated by this observation, we propose a novel test-time adaptation framework that leverages two self-supervised auxiliary tasks to help the primary forecasting network adapt to the test sequence. In the testing phase, our model can adjust the model parameters by several gradient updates to improve the generation quality. However, due to catastrophic forgetting, both auxiliary tasks typically tend to the low ability to automatically present the desired positive incentives for the final prediction performance. For this reason, we also propose a meta-auxiliary learning scheme for better adaptation. In terms of general setup, our approach obtains higher accuracy, and under two new experimental designs for out-of-distribution data (unseen subjects and categories), achieves significant improvements.
翻译:从历史观察序列中预测高保真度未来人体姿态,对于智能机器人与人类交互具有决定性作用。深度端到端学习方法——通常在外源数据集上训练通用预训练模型后直接应用于所有测试样本——已成为解决该问题的主流方案。尽管已取得令人鼓舞的进展,此类方法仍非最优方案,因为特定序列的独特属性(如运动风格、节奏)无法实现自适应调整。更普遍的情况是,在测试阶段,当遇到未见运动类别(分布外)时,预测姿态往往不可靠。受此启发,我们提出一种新颖的测试时自适应框架,通过两个自监督辅助任务协助主预测网络适应测试序列。在测试阶段,模型可通过数次梯度更新调整参数以提升生成质量。然而,由于灾难性遗忘问题,两个辅助任务通常难以自发对最终预测性能产生预期的正向激励。为此,我们进一步提出元辅助学习机制以实现更优自适应。在常规设定下,本方法可获得更高精度;针对分布外数据(未见主体与类别)的两类新实验设计,本方法均实现了显著性能提升。