Traditional visual servoing methods suffer from serving between scenes from multiple perspectives, which humans can complete with visual signals alone. In this paper, we investigated how multi-perspective visual servoing could be solved under robot-specific constraints, including self-collision, singularity problems. We presented a novel learning-based multi-perspective visual servoing framework, which iteratively estimates robot actions from latent space representations of visual states using reinforcement learning. Furthermore, our approaches were trained and validated in a Gazebo simulation environment with connection to OpenAI/Gym. Through simulation experiments, we showed that our method can successfully learn an optimal control policy given initial images from different perspectives, and it outperformed the Direct Visual Servoing algorithm with mean success rate of 97.0%.
翻译:传统视觉伺服方法在不同视角间的场景切换中表现不佳,而人类仅凭视觉信号即可完成此类任务。本文研究了在机器人特定约束(包括自碰撞和奇异性问题)下如何解决多视角视觉伺服问题。我们提出了一种新颖的基于学习的多视角视觉伺服框架,该框架利用强化学习从视觉状态的潜在空间表示中迭代估计机器人动作。此外,我们的方法在结合OpenAI/Gym的Gazebo仿真环境中进行了训练与验证。通过仿真实验证明,该方法在给定不同视角的初始图像时能够成功学习最优控制策略,并以97.0%的平均成功率优于直接视觉伺服算法。