Continuum robots are promising candidates for interactive tasks in various applications due to their unique shape, compliance, and miniaturization capability. Accurate and real-time shape sensing is essential for such tasks yet remains a challenge. Embedded shape sensing has high hardware complexity and cost, while vision-based methods require stereo setup and struggle to achieve real-time performance. This paper proposes the first eye-to-hand monocular approach to continuum robot shape sensing. Utilizing a deep encoder-decoder network, our method, MoSSNet, eliminates the computation cost of stereo matching and reduces requirements on sensing hardware. In particular, MoSSNet comprises an encoder and three parallel decoders to uncover spatial, length, and contour information from a single RGB image, and then obtains the 3D shape through curve fitting. A two-segment tendon-driven continuum robot is used for data collection and testing, demonstrating accurate (mean shape error of 0.91 mm, or 0.36% of robot length) and real-time (70 fps) shape sensing on real-world data. Additionally, the method is optimized end-to-end and does not require fiducial markers, manual segmentation, or camera calibration. Code and datasets will be made available at https://github.com/ContinuumRoboticsLab/MoSSNet.
翻译:连续型机器人因其独特的形状、柔顺性和微型化能力,在各种应用场景的交互任务中展现出巨大潜力。精确且实时的形状感知对此类任务至关重要,但仍是当前挑战。嵌入式形状传感硬件复杂度与成本高,而基于视觉的方法需立体设置且难以实现实时性能。本文提出首个面向连续型机器人形状感知的跟-到手单目方法。该方法采用深度编码器-解码器网络(MoSSNet),消除了立体匹配的计算开销并降低了对传感硬件的需求。具体而言,MoSSNet由一个编码器和三个并行解码器组成,从单一RGB图像中提取空间、长度和轮廓信息,并通过曲线拟合获取三维形状。我们采用两段肌腱驱动连续型机器人进行数据采集和测试,在真实数据上实现了精确(平均形状误差0.91毫米,即机器人长度的0.36%)且实时(70帧/秒)的形状感知。此外,该方法可端到端优化,无需基准标记、人工分割或相机标定。代码与数据集将在https://github.com/ContinuumRoboticsLab/MoSSNet上公开。