The rapid development of intelligent control methodologies has endowed robots with powerful autonomous intelligence. Cable routing, a ubiquitous foundational task in industry, provides a rigorous benchmark for robotic dexterity and sequential decision-making. In these practical scenarios, image observation distortion frequently occurs. Samples characterized by low-quality image observations often hinder accurate model training, posing challenges to the reliability and accuracy of intelligent control systems. Nevertheless, no dedicated intelligent control solution has been proposed for scenarios of image signal distortion. Meanwhile, image quality information has not been sufficiently exploited to further enhance the performance of intelligent control methodologies. To this end, we propose a novel robotic imitation learning framework that comprises an image quality assessment module, a confidence-based learning mechanism, and a decision-making module, which is designed to maintain high performance even under distorted image observations. In the proposed framework, the image quality assessment module synergizes with the confidence-based learning mechanism to enhance the efficacy of the decision-making module. Specifically, the image quality assessment module is incorporated to extract image quality information from image observations, while the confidence-based learning mechanism adaptively prioritizes challenging samples to improve learning effectiveness. The decision-making module determines appropriate discrete skills or continuous actions. Experimental results demonstrate that our formulated framework enhances the overall performance of the decision-making module.
翻译:智能控制方法的快速发展赋予了机器人强大的自主智能。线缆布线作为工业中一项普遍的基础任务,为机器人的灵巧操作和序列决策能力提供了严格的基准。在实际应用场景中,图像观测畸变频繁发生。以低质量图像观测为特征的样本常常阻碍模型的精确训练,对智能控制系统的可靠性和准确性构成挑战。然而,目前尚无针对图像信号畸变场景的专门智能控制解决方案。与此同时,图像质量信息也未得到充分挖掘以进一步提升智能控制方法的性能。为此,我们提出了一种新颖的机器人模仿学习框架,该框架包含图像质量评估模块、基于置信度的学习机制和决策模块,旨在即使在畸变图像观测下也能保持高性能。在所提出的框架中,图像质量评估模块与基于置信度的学习机制协同作用,以增强决策模块的效能。具体而言,引入图像质量评估模块用于从图像观测中提取图像质量信息,而基于置信度的学习机制则自适应地优先处理困难样本以提高学习效果。决策模块可确定合适的离散技能或连续动作。实验结果表明,我们提出的框架提升了决策模块的整体性能。