Humanoid robots hold immense potential for real-world assistance, yet agile interaction with objects in unstructured environments demands tightly coupled whole-body coordination. Despite recent advancements, current controllers face a critical deployment gap. They rely heavily on dense reference trajectories and perfect state observability, which inherently limits physical generalization. We present Vision Guided Agile Interaction Control (VAIC), a unified framework that bridges this gap by operating exclusively on onboard depth, historical proprioception, and a decoupled user command interface. VAIC employs a two-stage distillation paradigm. First, a privileged teacher policy masters diverse interaction skills using precise object kinematics and exact environmental states. Second, a deployable student policy distills these capabilities by replacing full body tracking with velocity targets across multiple axes and an interaction indicator for each frame. The student utilizes a recurrent object adaptation module to implicitly infer unobservable object dynamics from raw depth streams and proprioception. Evaluations and real-world deployments on the humanoid robot demonstrate that a single VAIC policy successfully executes highly diverse dynamic tasks. These tasks include box carrying, cart interaction, and skateboarding, consistently outperforming baselines and advancing autonomous humanoid deployment.
翻译:仿人机器人在现实世界辅助中具有巨大潜力,然而在非结构化环境中与物体进行敏捷交互需要紧密耦合的全身协调。尽管近期取得了进展,现有控制器仍面临关键部署瓶颈。它们严重依赖密集的参考轨迹与完美状态可观测性,这从根本上限制了物理泛化能力。我们提出视觉引导的敏捷交互控制(VAIC),这是一个统一框架,通过仅依赖机载深度数据、历史本体感知和解耦的用户指令接口来弥补这一不足。VAIC采用两阶段蒸馏范式。首先,特权教师策略利用精确的物体运动学和环境状态掌握多样化交互技能。其次,可部署学生策略通过将全身跟踪替换为多轴速度目标及每帧交互指示符来精简这些能力。学生策略利用循环物体自适应模块,从原始深度流和本体感知中隐式推断不可观测的物体动力学。在仿人机器人上的评估和实际部署表明,单一VAIC策略能成功执行高度多样化的动态任务,包括搬箱、推车和滑板,持续优于基线方法,推动自主仿人机器人部署的发展。