The rapid development of collaborative robotics has provided a new possibility of helping the elderly who has difficulties in daily life, allowing robots to operate according to specific intentions. However, efficient human-robot cooperation requires natural, accurate and reliable intention recognition in shared environments. The current paramount challenge for this is reducing the uncertainty of multimodal fused intention to be recognized and reasoning adaptively a more reliable result despite current interactive condition. In this work we propose a novel learning-based multimodal fusion framework Batch Multimodal Confidence Learning for Opinion Pool (BMCLOP). Our approach combines Bayesian multimodal fusion method and batch confidence learning algorithm to improve accuracy, uncertainty reduction and success rate given the interactive condition. In particular, the generic and practical multimodal intention recognition framework can be easily extended further. Our desired assistive scenarios consider three modalities gestures, speech and gaze, all of which produce categorical distributions over all the finite intentions. The proposed method is validated with a six-DoF robot through extensive experiments and exhibits high performance compared to baselines.
翻译:协作机器人的快速发展为帮助日常生活困难的老年人提供了新的可能性,使机器人能够根据特定意图进行操作。然而,高效的人机协作需要在共享环境中实现自然、准确且可靠的意图识别。当前面临的核心挑战在于降低待识别多模态融合意图的不确定性,并能够根据当前交互条件自适应地推理出更可靠的结果。本文提出了一种新颖的基于学习的多模态融合框架——面向意见池的批量多模态置信度学习(BMCLOP)。该方法结合了贝叶斯多模态融合方法与批量置信度学习算法,以在给定交互条件下提高识别准确性、降低不确定性并提升成功率。特别地,该通用且实用的多模态意图识别框架易于进一步扩展。我们设计的辅助场景考虑了手势、语音和视线三种模态,所有模态均针对有限意图集合生成类别分布。所提方法通过大量实验在六自由度机器人上进行了验证,与基线方法相比展现了优越的性能。