In this work, we first formulate the problem of robotic water scooping using goal-conditioned reinforcement learning. This task is particularly challenging due to the complex dynamics of fluids and the need to achieve multi-modal goals. The policy is required to successfully reach both position goals and water amount goals, which leads to a large convoluted goal state space. To overcome these challenges, we introduce Goal Sampling Adaptation for Scooping (GOATS), a curriculum reinforcement learning method that can learn an effective and generalizable policy for robot scooping tasks. Specifically, we use a goal-factorized reward formulation and interpolate position goal distributions and amount goal distributions to create curriculum throughout the learning process. As a result, our proposed method can outperform the baselines in simulation and achieves 5.46% and 8.71% amount errors on bowl scooping and bucket scooping tasks, respectively, under 1000 variations of initial water states in the tank and a large goal state space. Besides being effective in simulation environments, our method can efficiently adapt to noisy real-robot water-scooping scenarios with diverse physical configurations and unseen settings, demonstrating superior efficacy and generalizability. The videos of this work are available on our project page: https://sites.google.com/view/goatscooping.
翻译:摘要:本研究首先利用目标条件强化学习形式化了机器人舀水问题。由于流体动力学的复杂性和多模态目标的实现需求,该任务极具挑战性。策略需同时达成位置目标与水量目标,导致目标状态空间高度复杂且相互交织。为应对这些挑战,我们提出舀水任务目标采样自适应方法(GOATS),这是一种课程强化学习方法,能够为机器人舀水任务学习有效且可泛化的策略。具体而言,我们采用基于目标因子化的奖励函数,通过在学习过程中对位置目标分布与水量目标分布进行插值构建课程。实验结果表明,所提方法在仿真环境中优于基线方法,在水箱初始状态包含1000种变化且目标状态空间庞大的条件下,碗状舀水与桶状舀水任务的水量误差分别仅为5.46%和8.71%。除在仿真环境中表现优异外,该方法还能高效适应具有多种物理配置和未知设置的含噪真实机器人舀水场景,展现出卓越的有效性与泛化能力。本工作的视频演示可访问项目页面:https://sites.google.com/view/goatscooping。