We explore multi-log grasping using reinforcement learning and virtual visual servoing for automated forwarding in a simulated environment. Automation of forest processes is a major challenge, and many techniques regarding robot control pose different challenges due to the unstructured and harsh outdoor environment. Grasping multiple logs involves various problems of dynamics and path planning, where understanding the interaction between the grapple, logs, terrain, and obstacles requires visual information. To address these challenges, we separate image segmentation from crane control and utilise a virtual camera to provide an image stream from reconstructed 3D data. We use Cartesian control to simplify domain transfer to real-world applications. Since log piles are static, visual servoing using a 3D reconstruction of the pile and its surroundings is equivalent to using real camera data until the point of grasping. This relaxes the limits on computational resources and time for the challenge of image segmentation and allows for collecting data in situations where the log piles are not occluded. The disadvantage is the lack of information during grasping. We demonstrate that this problem is manageable and present an agent that is 95% successful in picking one or several logs from challenging piles of 2--5 logs.
翻译:我们探索了在模拟环境中利用强化学习和虚拟视觉伺服技术实现多根原木自主抓取的自动化运输方案。森林作业自动化是重大挑战,由于户外环境非结构化且条件严苛,许多机器人控制技术面临不同难题。多根原木抓取涉及动力学与路径规划的多重问题,理解抓取臂、原木、地形和障碍物之间的相互作用需要视觉信息。为应对这些挑战,我们将图像分割与起重机控制分离,利用虚拟相机从重建的三维数据中获取图像流。采用笛卡尔控制以简化向实际应用的域迁移。由于原木堆静态分布,基于三维重建的视觉伺服技术在抓取前可等效使用真实相机数据,这降低了图像分割对计算资源和时间的要求,并允许在无遮挡场景下收集数据。其劣势在于抓取过程中信息缺失。我们证明该问题具有可控性,并展示了在2-5根原木组成的复杂堆叠中,智能体抓取单根或多根原木的成功率达到95%。