Despite recent advances in control, reinforcement learning, and imitation learning, developing a unified framework that can achieve agile, precise, and robust whole-body behaviors, particularly in long-horizon tasks, remains challenging. Existing approaches typically follow two paradigms: coupled whole-body policies for global coordination and decoupled policies for modular precision. However, without a systematic method to integrate both, this trade-off between agility, robustness, and precision remains unresolved. In this work, we propose BAT, an online policy-switching framework that dynamically selects between two complementary whole-body RL controllers to balance agility and stability across different motion contexts. Our framework consists of two complementary modules: a switching policy learned via hierarchical RL with an expert guidance from sliding-horizon policy pre-evaluation, and an option-aware VQ-VAE that predicts option preference from discrete motion token sequences for improved generalization. The final decision is obtained via confidence-weighted fusion of two modules. Extensive simulations and real-world experiments on the Unitree G1 humanoid robot demonstrate that BAT enables versatile long-horizon loco-manipulation and outperforms prior methods across diverse tasks.


翻译:尽管控制、强化学习和模仿学习领域已取得显著进展,但开发能够实现敏捷、精确且鲁棒的全身行为(尤其在长时域任务中)的统一框架仍具挑战性。现有方法通常遵循两种范式:用于全局协调的耦合全身策略与用于模块化精确性的解耦策略。然而,由于缺乏系统化的融合方法,敏捷性、鲁棒性与精确性之间的权衡问题仍未解决。本文提出BAT——一种在线策略切换框架,通过动态选择两个互补的全身强化学习控制器,在不同运动场景中平衡敏捷性与稳定性。该框架包含两个互补模块:基于分层强化学习与滑动时域策略预评估专家引导训练的切换策略,以及通过离散运动标记序列预测选项偏好以提升泛化能力的选项感知VQ-VAE。最终决策通过两个模块的置信度加权融合生成。在宇树G1人形机器人上的大规模仿真与真实世界实验表明,BAT能够实现多样化的长时域移动操作任务,并在各类任务中均优于既有方法。

0
下载
关闭预览

相关内容

BAT,分别指21世纪10年代,中国大陆互联网的三大巨头:百度(Baidu),阿里巴巴(Alibaba),腾讯(Tencent)
国外有人/无人平台协同作战概述
无人机
124+阅读 · 2019年5月28日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
24+阅读 · 2011年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
VIP会员
最新内容
非对称防御中的自组织临界性:俄乌战争
专知会员服务
7+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
6+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
11+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
12+阅读 · 8月8日
相关VIP内容
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
24+阅读 · 2011年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员