Enabling humanoid robots to follow free-form natural language commands is a critical step toward seamless human-robot interaction and general-purpose embodied AI. However, existing methods remain limited, often constrained to simple instructions or forced to sacrifice motion diversity for physical plausibility. To address this gap, we present Humanoid-LLA, a Large Language Action model that translates unconstrained natural language directly into executable whole-body motions for humanoid robots. Our approach tackles two core challenges: paired language-humanoid motion data scarcity and physical instability. First, we bridge high-level language semantics with physically-grounded control by learning a unified human-humanoid motion vocabulary. Second, we introduce a novel two-stage fine-tuning framework that begins with supervised motion Chain-of-Thought learning, followed by reinforcement learning refined with physical feedback to ensure robustness and stability. Extensive evaluation in simulation and real-world cross-embodiment experiments demonstrates that Humanoid-LLA achieves superior generalization to novel language commands and diverse motion generation while maintaining high physical fidelity.
翻译:实现人形机器人遵循自由形式的自然语言指令,是人机无缝交互及通用具身智能的关键一步。然而,现有方法仍存在局限,常受限于简单指令,或被迫牺牲动作多样性以换取物理合理性。为弥补这一不足,我们提出Humanoid-LLA——一种大规模语言动作模型,能将未受约束的自然语言直接转化为可执行的人形机器人全身动作。我们的方法应对两大核心挑战:语言与人形机器人配对动作数据稀缺及物理不稳定性。首先,通过学习统一的人-人形机器人运动词汇,我们桥接了高层语言语义与物理层面控制。其次,我们引入一种新颖的两阶段微调框架:先进行监督式动作思维链学习,随后结合物理反馈进行强化学习精调,以确保鲁棒性和稳定性。在仿真及真实世界跨本体实验中的广泛评估表明,Humanoid-LLA在应对新奇语言指令和多样化动作生成方面实现了卓越的泛化能力,同时保持了高物理保真度。