Vision-language-action (VLA) models provide a promising paradigm for scalable robotic manipulation, yet their reliance on success-only behavioral cloning leaves them brittle; lacking corrective training signals, minor execution errors rapidly compound into unrecoverable, out-of-distribution failures. To address this limitation, we propose Adaptive Failure-Informed Learning (AFIL), an end-to-end framework that leverages failure trajectories as adaptive negative guidance for diffusion- and flow-based VLA policies. AFIL uses a pretrained VLA to generate failure rollouts online, avoiding the need for handcrafted failure-mode design or human-in-the-loop recovery. It then jointly trains Dual Action Generators (DAGs) for successful and failed behaviors while sharing a common vision-language backbone, enabling efficient failure-aware policy learning with limited parameter overhead. During sampling, the failure generator adaptively steers action generation away from failure-prone regions and toward more reliable success modes, with guidance strength determined by the per-diffusion-step distance between success and failure distributions. Experiments across in-domain and out-of-domain robotic manipulation tasks, covering both short- and long-horizon settings, show that AFIL consistently improves task success rates and robustness over existing VLA baselines, demonstrating its effectiveness, efficiency, and generality.
翻译:视觉-语言-动作(VLA)模型为可扩展的机器人操作提供了有前景的范式,但其仅依赖成功示例的行为克隆使其脆弱不堪;由于缺乏纠错训练信号,微小的执行错误会迅速累积成不可恢复的、超出分布范围的失败。为解决这一局限,我们提出自适应失败信息学习(AFIL)——一种端到端框架,利用失败轨迹作为扩散和流式VLA策略的自适应负向引导。AFIL使用预训练的VLA在线生成失败滚动轨迹,避免了手工设计失败模式或人工干预恢复的需求。随后,它在共享共同视觉-语言骨干网络的同时,联合训练针对成功和失败行为的双动作生成器(DAGs),从而以有限的参数开销实现高效的失败感知策略学习。在采样过程中,失败生成器自适应地将动作生成引导远离易失败区域,转向更可靠的成功模式,其引导强度由成功与失败分布间逐扩散步骤的距离决定。涵盖短程和长程场景的域内与域外机器人操作任务实验表明,AFIL相较于现有VLA基线模型,始终能提升任务成功率和鲁棒性,验证了其有效性、高效性和泛化能力。