While sim2real efforts are necessary for effective policy transfer to hardware, there is such a thing as too much of a good thing. We argue that sim2real efforts have led to misaligned incentives with policy learning, resulting in simulator lock in and poor policy exploration due to the unreasonable constraints imposed by the real world. We offer a diagnosis and explanation of the current status of the problem, and propose a potential solution via a sim2sim2real paradigm that leverages the robot's kinematics as the sole design constraint.
翻译:虽然 sim2real 努力对于策略向硬件的有效迁移是必要的,但“过犹不及”。我们认为,sim2real 努力已导致与策略学习激励不一致,由于真实世界施加的不合理约束,造成模拟器锁定和策略探索不足。我们对问题的现状进行了诊断和解释,并提出了一个潜在的解决方案:通过利用机器人运动学作为唯一设计约束的 sim2sim2real 范式。