Dexterous in-hand manipulation for a multi-fingered anthropomorphic hand is extremely difficult because of the high-dimensional state and action spaces, rich contact patterns between the fingers and objects. Even though deep reinforcement learning has made moderate progress and demonstrated its strong potential for manipulation, it is still faced with certain challenges, such as large-scale data collection and high sample complexity. Especially, for some slight change scenes, it always needs to re-collect vast amounts of data and carry out numerous iterations of fine-tuning. Remarkably, humans can quickly transfer learned manipulation skills to different scenarios with little supervision. Inspired by human flexible transfer learning capability, we propose a novel dexterous in-hand manipulation progressive transfer learning framework (PTL) based on efficiently utilizing the collected trajectories and the source-trained dynamics model. This framework adopts progressive neural networks for dynamics model transfer learning on samples selected by a new samples selection method based on dynamics properties, rewards and scores of the trajectories. Experimental results on contact-rich anthropomorphic hand manipulation tasks show that our method can efficiently and effectively learn in-hand manipulation skills with a few online attempts and adjustment learning under the new scene. Compared to learning from scratch, our method can reduce training time costs by 95%.
翻译:多指拟人手进行灵巧手内操作极具挑战性,原因在于高维状态与动作空间、手指与物体间丰富的接触模式。尽管深度强化学习已取得一定进展并展现出强大的操作潜力,但仍面临大规模数据收集和高样本复杂度等挑战。特别是对于某些细微变化的场景,往往需要重新收集大量数据并进行大量迭代微调。值得注意的是,人类能够快速将已学习的操作技能迁移至不同场景,且只需极少监督。受人类灵活迁移学习能力的启发,我们提出了一种新颖的灵巧手内操作渐进式迁移学习框架(PTL),该框架基于高效利用收集的轨迹和源训练动力学模型。该框架采用渐进式神经网络,根据一种新的基于动力学特性、轨迹奖励与得分的样本选择方法所筛选的样本,进行动力学模型迁移学习。在富含接触的拟人手操作任务上的实验结果表明,我们的方法能够在新场景下通过少量在线尝试和调整学习,高效且有效地学习手内操作技能。与从头学习相比,我们的方法可将训练时间成本降低95%。