As human-robot collaboration increases in the workforce, it becomes essential for human-robot teams to coordinate efficiently and intuitively. Traditional approaches for human-robot scheduling either utilize exact methods that are intractable for large-scale problems and struggle to account for stochastic, time varying human task performance, or application-specific heuristics that require expert domain knowledge to develop. We propose a deep learning-based framework, called HybridNet, combining a heterogeneous graph-based encoder with a recurrent schedule propagator for scheduling stochastic human-robot teams under upper- and lower-bound temporal constraints. The HybridNet's encoder leverages Heterogeneous Graph Attention Networks to model the initial environment and team dynamics while accounting for the constraints. By formulating task scheduling as a sequential decision-making process, the HybridNet's recurrent neural schedule propagator leverages Long Short-Term Memory (LSTM) models to propagate forward consequences of actions to carry out fast schedule generation, removing the need to interact with the environment between every task-agent pair selection. The resulting scheduling policy network provides a computationally lightweight yet highly expressive model that is end-to-end trainable via Reinforcement Learning algorithms. We develop a virtual task scheduling environment for mixed human-robot teams in a multi-round setting, capable of modeling the stochastic learning behaviors of human workers. Experimental results showed that HybridNet outperformed other human-robot scheduling solutions across problem sizes for both deterministic and stochastic human performance, with faster runtime compared to pure-GNN-based schedulers.
翻译:随着人机协作在劳动力中的增加,人机团队高效且直观的协调变得至关重要。传统的人机调度方法要么采用精确方法,但难以处理大规模问题且难以应对随机时变的人类任务绩效,要么采用依赖专家领域知识开发的特定应用启发式方法。我们提出了一种名为HybridNet的深度学习框架,该框架结合了基于异构图的编码器和循环调度传播器,用于在上界和下界时间约束下调度随机人机团队。HybridNet的编码器利用异构图表征网络对初始环境和团队动态进行建模,同时兼顾约束条件。通过将任务调度表述为序贯决策过程,HybridNet的循环神经调度传播器利用长短期记忆模型(LSTM)向前传播动作后果以快速生成调度方案,从而无需在每次任务-智能体配对选择之间与环境交互。由此产生的调度策略网络提供了一种计算轻量级但高度表达性的模型,可通过强化学习算法进行端到端训练。我们开发了一个用于多轮混合人机团队的虚拟任务调度环境,该环境能够对工人随机的学习行为进行建模。实验结果表明,HybridNet在不同问题规模下(针对确定性和随机性人类绩效情况)均优于其他人机调度方案,且与纯图神经网络(GNN)基础调度器相比,具有更快的运行时间。