In this work, we first formulate the problem of goal-conditioned robotic water scooping with reinforcement learning. This task is challenging due to the complex dynamics of fluid and multi-modal goal-reaching. The policy is required to achieve both position goals and water amount goals, which leads to a large convoluted goal state space. To address these challenges, we introduce Goal Sampling Adaptation for Scooping (GOATS), a curriculum reinforcement learning method that can learn an effective and generalizable policy for robot scooping tasks. Specifically, we use a goal-factorized reward formulation and interpolate position goal distributions and amount goal distributions to create curriculum through the learning process. As a result, our proposed method can outperform the baselines in simulation and achieves 5.46% and 8.71% amount errors on bowl scooping and bucket scooping tasks, respectively, under 1000 variations of initial water states in the tank and a large goal state space. Besides being effective in simulation environments, our method can efficiently generalize to noisy real-robot water-scooping scenarios with different physical configurations and unseen settings, demonstrating superior efficacy and generalizability. The videos of this work are available on our project page: https://sites.google.com/view/goatscooping.
翻译:本文首先将目标条件机器人铲水问题形式化为强化学习任务。由于流体动力学复杂性及多模态目标达成特性,该任务具有显著挑战性。策略需同时实现位置目标与水量目标,导致目标状态空间高度耦合且维度庞大。为应对这些挑战,我们提出铲取目标采样自适应方法(GOATS),这是一种课程强化学习方法,能够为机器人铲取任务学习有效且可泛化的策略。具体而言,我们采用目标分解式奖励函数设计,通过插值位置目标分布与水量目标分布,在训练过程中构建课程学习路径。实验结果表明,在包含1000种初始水箱状态变化及大规模目标状态空间的设定下,本方法在碗铲取与桶铲取任务中分别实现了5.46%和8.71%的水量误差,性能优于基线方法。除仿真环境有效性外,该方法还能高效泛化至具有不同物理构型与未知场景的真实机器人铲水噪声环境,展现出卓越的效能与泛化能力。本研究演示视频可参见项目主页:https://sites.google.com/view/goatscooping。