Machine scheduling aims to optimize job assignments to machines while adhering to manufacturing rules and job specifications. This optimization leads to reduced operational costs, improved customer demand fulfillment, and enhanced production efficiency. However, machine scheduling remains a challenging combinatorial problem due to its NP-hard nature. Deep Reinforcement Learning (DRL), a key component of artificial general intelligence, has shown promise in various domains like gaming and robotics. Researchers have explored applying DRL to machine scheduling problems since 1995. This paper offers a comprehensive review and comparison of DRL-based approaches, highlighting their methodology, applications, advantages, and limitations. It categorizes these approaches based on computational components: conventional neural networks, encoder-decoder architectures, graph neural networks, and metaheuristic algorithms. Our review concludes that DRL-based methods outperform exact solvers, heuristics, and tabular reinforcement learning algorithms in terms of computation speed and generating near-global optimal solutions. These DRL-based approaches have been successfully applied to static and dynamic scheduling across diverse machine environments and job characteristics. However, DRL-based schedulers face limitations in handling complex operational constraints, configurable multi-objective optimization, generalization, scalability, interpretability, and robustness. Addressing these challenges will be a crucial focus for future research in this field. This paper serves as a valuable resource for researchers to assess the current state of DRL-based machine scheduling and identify research gaps. It also aids experts and practitioners in selecting the appropriate DRL approach for production scheduling.
翻译:机器调度旨在遵循制造规则和作业规格的前提下优化作业与机器的分配,从而降低运营成本、提升客户需求满足度并提高生产效率。然而,由于机器调度属于NP-hard组合优化问题,其求解仍具挑战性。深度强化学习(DRL)作为通用人工智能的关键组成部分,已在游戏、机器人等领域展现出巨大潜力。自1995年起,研究者开始探索将DRL应用于机器调度问题。本文系统综述并比较了基于DRL的调度方法,重点阐述其方法论、应用场景、优势与局限性。我们将这些方法按计算组件进行分类:传统神经网络、编码器-解码器架构、图神经网络以及元启发式算法。综述表明,基于DRL的方法在计算速度和生成近全局最优解方面优于精确求解器、启发式算法和表格型强化学习算法。这些方法已成功应用于不同机器环境与作业特征的静态和动态调度场景。然而,基于DRL的调度器在处理复杂操作约束、可配置多目标优化、泛化性、可扩展性、可解释性及鲁棒性方面仍存在局限。解决这些挑战将是该领域未来研究的重点方向。本文可为研究者评估基于DRL的机器调度现状、发现研究空白提供参考,同时帮助专家和实践者选择适合生产调度的DRL方法。