Humanoid-Object Interaction (HOI) is a fundamental capability for humanoid robots, yet it remains challenging due to the tight coupling between dynamic balance and stable interaction with diverse objects. Existing methods often require time-consuming task-specific policy training or rely on rigid trajectory replay, which limits their ability to accommodate novel interaction scenarios. In this work, we present \textit{GenHOI}, a simple yet effective framework that enables humanoid robots to perform diverse object-interaction tasks in a zero-shot manner by directly imitating a single generated video, without task-specific training or physical demonstration data. GenHOI first reconstructs the robot-object scene in simulation and renders a first-frame image, which, together with the language command, conditions the synthesis of a task-oriented interaction video. The generated video is then analyzed to identify interaction-relevant contact events and estimate hand-object contact regions, which are encoded as object-centric geometric constraints that convert visual interaction cues into physically grounded optimization priors. Guided by these priors, the reference motion recovered from the video is refined and smoothed to resolve the scale ambiguity inherent in 2D video generation, while adapting a single reference trajectory to unseen robot-object relative poses. The optimized trajectory is finally executed by a closed-loop tracking controller. We validate the proposed framework in extensive simulation and real-world experiments across diverse object-interaction tasks, including box grasping, asymmetric bimanual chair carrying, table lifting from below, and cylindrical-object enveloping.
翻译:人形物体交互(HOI)是人形机器人的一项基础能力,但由于动态平衡与多样化物体稳定交互之间的紧密耦合,该任务仍然具有挑战性。现有方法通常需要耗时的任务特定策略训练,或依赖刚性的轨迹回放,这限制了其适应新型交互场景的能力。本文提出GenHOI——一种简洁而高效的框架,使人形机器人能够通过直接模仿单个生成视频,在零样本条件下执行多样化的物体交互任务,无需任务特定训练或物理示教数据。GenHOI首先在仿真中重建机器人-物体场景并渲染首帧图像,该图像与语言指令共同约束任务导向交互视频的合成。随后对生成视频进行分析,以识别与交互相关的接触事件,并估计手-物体接触区域,这些区域被编码为以物体为中心的几何约束,将视觉交互线索转化为物理驱动的优化先验。在这些先验的引导下,从视频中恢复的参考运动被精炼和平滑化,以解决二维视频生成中固有的尺度模糊性,并同时使单一参考轨迹适应未见过的机器人-物体相对位姿。最终通过闭环跟踪控制器执行优化后的轨迹。我们在涵盖多种物体交互任务(包括抓取箱子、非对称双臂搬运椅子、自下而上抬举桌子以及抱持圆柱形物体)的大量仿真与真实世界实验中验证了所提框架的有效性。