Offering great potential in robotic manipulation, a capable Vision-Language-Action (VLA) foundation model is expected to faithfully generalize across tasks and platforms while ensuring cost efficiency (e.g., data and GPU hours required for adaptation). To this end, we develop LingBot-VLA with around 20,000 hours of real-world data from 9 popular dual-arm robot configurations. Through a systematic assessment on 4 robotic platforms, each completing 100 tasks with 130 post-training episodes per task, our model achieves clear superiority over competitors, showcasing its strong performance and broad generalizability. We have also built an efficient codebase, which delivers a throughput of 261 samples per second with an 8-GPU training setup, representing a 1.5~2.8$\times$ (depending on the relied VLM base model) speedup over existing VLA-oriented codebases. The above features ensure that our model is well-suited for real-world deployment. To advance the field of robot learning, we provide open access to the code, base model, and benchmark data, with a focus on enabling more challenging tasks and promoting sound evaluation standards.
翻译:在机器人操作领域,一个强大的视觉-语言-动作(VLA)基础模型有望在跨任务和跨平台上实现可靠的泛化,同时确保成本效益(例如,适应所需的数据和GPU时间)。为此,我们基于来自9种主流双臂机器人配置的约20,000小时真实世界数据,开发了LingBot-VLA。通过在4个机器人平台上进行系统评估,每个平台完成100个任务,每个任务包含130次后训练回合,我们的模型相较于竞争对手展现出明显的优势,证明了其强大的性能和广泛的泛化能力。我们还构建了一个高效的代码库,在8-GPU训练设置下可实现每秒261个样本的吞吐量,相比现有的VLA导向代码库实现了1.5~2.8倍(取决于所依赖的VLM基础模型)的加速。上述特性确保我们的模型非常适用于实际部署。为推动机器人学习领域的发展,我们公开提供代码、基础模型和基准数据,重点关注更具挑战性的任务和促进合理的评估标准。