Imitation learning (IL) policies in robotics deliver strong performance in controlled settings but remain brittle in real-world deployments: rare events such as hardware faults, defective parts, unexpected human actions, or any state that lies outside the training distribution can lead to failed executions. Vision-based Anomaly Detection (AD) methods emerged as an appropriate solution to detect these anomalous failure states but do not distinguish failures from benign deviations. We introduce FIDeL (Failure Identification in Demonstration Learning), a policy-independent failure detection module. Leveraging recent AD methods, FIDeL builds a compact representation of nominal demonstrations and aligns incoming observations via optimal transport matching to produce anomaly scores and heatmaps. Spatio-temporal thresholds are derived with an extension of conformal prediction, and a Vision-Language Model (VLM) performs semantic filtering to discriminate benign anomalies from genuine failures. We also introduce BotFails, a multimodal dataset of real-world tasks for failure detection in robotics. FIDeL consistently outperforms state-of-the-art baselines, yielding +5.30% percent AUROC in anomaly detection and +17.38% percent failure-detection accuracy on BotFails compared to existing methods.
翻译:模仿学习(IL)策略在受控环境中表现出色,但在实际部署中仍显脆弱:硬件故障、部件缺陷、意外人为动作等罕见事件,或任何超出训练分布的状态,均可能导致执行失败。基于视觉的异常检测(AD)方法可作为检测此类异常故障状态的解决方案,但无法区分故障与非良性偏差。我们提出FIDeL(演示学习中的故障识别模块),一种策略无关的故障检测模块。该模块利用现有AD方法构建标称示范的紧凑表征,并通过最优传输匹配对齐实时观测以生成异常分数与热力图。通过扩展共形预测方法推导时空阈值,结合视觉语言模型(VLM)进行语义滤波,区分良性偏差与真实故障。我们还引入BotFails多模态数据集,专用于机器人故障检测的实际任务。与现有方法相比,FIDeL在BotFails数据集上持续超越最优基准,异常检测AUROC提升5.30%,故障检测准确率提升17.38%。