We address language-conditioned robotic manipulation using flow-based trajectory generation, which enables training on human and web videos of object manipulation and requires only minimal embodiment-specific data. This task is challenging, as object trajectory generation from pre-manipulation images and natural language instructions requires appropriate instruction-flow alignment. To tackle this challenge, we propose the flow-based Language Instruction-guided open-Loop ACtion generator (LILAC). This flow-based Vision-Language-Action model (VLA) generates object-centric 2D optical flow from an RGB image and a natural language instruction, and converts the flow into a 6-DoF manipulator trajectory. LILAC incorporates two key components: Semantic Alignment Loss, which strengthens language conditioning to generate instruction-aligned optical flow, and Prompt-Conditioned Cross-Modal Adapter, which aligns learned visual prompts with image and text features to provide rich cues for flow generation. Experimentally, our method outperformed existing approaches in generated flow quality across multiple benchmarks. Furthermore, in physical object manipulation experiments using free-form instructions, LILAC demonstrated a superior task success rate compared to existing methods. The project page is available at https://lilac-75srg.kinsta.page/.
翻译:摘要:我们研究利用基于光流的轨迹生成实现语言条件下的机器人操作,该方法可通过人类操作物体的视频及网络视频进行训练,且仅需极少量具身特定数据。该任务具有挑战性,因为从操作前图像和自然语言指令生成物体轨迹需要恰当的指令-光流对齐。为应对这一挑战,我们提出基于光流的语言指令引导开环动作生成器(LILAC)。该基于光流的视觉-语言-动作模型(VLA)可从RGB图像与自然语言指令生成以物体为中心的二维光流,并将该光流转换为六自由度机械臂轨迹。LILAC包含两个关键组件:语义对齐损失(增强语言条件作用以生成与指令对齐的光流),以及提示条件跨模态适配器(将学习到的视觉提示与图像及文本特征对齐,为光流生成提供丰富线索)。实验表明,我们的方法在多个基准测试中生成的光流质量均优于现有方法。此外,在使用自由形式指令的物理物体操作实验中,LILAC的任务成功率显著高于现有方法。项目主页:https://lilac-75srg.kinsta.page/。