Recent advancements in model-free deep reinforcement learning have enabled efficient agent training. However, challenges arise when determining the region of attraction for these controllers, especially if the region does not fully cover the desired area. This paper addresses this issue by introducing a feedback motion control algorithm that utilizes data-driven techniques and neural networks. The algorithm constructs a graph of connected reinforcement-learning based controllers, each with its own defined region of attraction. This incremental approach effectively covers a bounded region of interest, creating a trajectory of interconnected nodes that guide the system from an initial state to the goal. Two approaches are presented for connecting nodes within the algorithm. The first is a tree-structured method, facilitating "point-to-point" control by constructing a tree connecting the initial state to the goal state. The second is a graph-structured method, enabling "space-to-space" control by building a graph within a bounded region. This approach allows for control from arbitrary initial and goal states. The proposed method's performance is evaluated on a first-order dynamic system, considering scenarios both with and without obstacles. The results demonstrate the effectiveness of the proposed algorithm in achieving the desired control objectives.
翻译:近年来,无模型深度强化学习的进展使得智能体训练变得高效。然而,在确定这些控制器的吸引域时仍面临挑战,尤其是当吸引域未完全覆盖目标区域时。本文针对这一问题,提出了一种融合数据驱动技术与神经网络的反馈运动控制算法。该算法通过构建一个由多个基于强化学习的控制器组成的连接图,每个控制器均定义其专属吸引域。这种增量式方法可有效覆盖有界目标区域,生成由互连节点构成的轨迹,引导系统从初始状态到达目标状态。算法中提出了两种节点连接方式:第一种是树形结构方法,通过构建连接初始状态与目标状态的树形结构实现"点对点"控制;第二种是图结构方法,通过在有界区域内构建图结构实现"空间到空间"控制,支持从任意初始状态到任意目标状态的控制。所提方法在一阶动态系统上进行性能评估,涵盖有障碍物与无障碍物两种场景。实验结果表明,该算法能有效实现预期控制目标。