Recently, there has been tremendous interest in industry 4.0 infrastructure to address labor shortages in global supply chains. Deploying artificial intelligence-enabled robotic bin picking systems in real world has become particularly important for reducing stress and physical demands of workers while increasing speed and efficiency of warehouses. To this end, artificial intelligence-enabled robotic bin picking systems may be used to automate order picking, but with the risk of causing expensive damage during an abnormal event such as sensor failure. As such, reliability becomes a critical factor for translating artificial intelligence research to real world applications and products. In this paper, we propose a reliable object detection and segmentation system with MultiModal Redundancy (MMRNet) for tackling object detection and segmentation for robotic bin picking using data from different modalities. This is the first system that introduces the concept of multimodal redundancy to address sensor failure issues during deployment. In particular, we realize the multimodal redundancy framework with a gate fusion module and dynamic ensemble learning. Finally, we present a new label-free multi-modal consistency (MC) score that utilizes the output from all modalities to measure the overall system output reliability and uncertainty. Through experiments, we demonstrate that in an event of missing modality, our system provides a much more reliable performance compared to baseline models. We also demonstrate that our MC score is a more reliability indicator for outputs during inference time compared to the model generated confidence scores that are often over-confident.
翻译:近年来,工业4.0基础设施在解决全球供应链劳动力短缺问题方面引起了极大关注。将人工智能赋能的机器人料箱抓取系统部署到实际场景中,对于降低工人的压力和体力需求、同时提高仓库的速度和效率尤为重要。为此,这类系统可用于自动化订单拣选,但存在因传感器故障等异常事件而导致昂贵损失的风险。因此,可靠性成为将人工智能研究转化为实际应用和产品的关键因素。本文提出了一种基于多模态冗余的可靠目标检测与分割系统(MMRNet),用于利用不同模态的数据解决机器人料箱抓取中的目标检测与分割任务。这是首个引入多模态冗余概念以应对部署过程中传感器故障问题的系统。具体而言,我们通过门控融合模块和动态集成学习实现了多模态冗余框架。最后,我们提出了一种新的无标签多模态一致性(MC)分数,该分数利用所有模态的输出衡量整个系统的输出可靠性和不确定性。实验表明,在模态缺失的情况下,我们的系统相比基线模型提供了更可靠的性能。我们还证明,与通常过于自信的模型生成置信度分数相比,我们的MC分数是推理过程中输出更可靠的指标。