The following interdisciplinary article presents a memetic algorithm with applying deep reinforcement learning (DRL) for solving practically oriented dual resource constrained flexible job shop scheduling problems (DRC-FJSSP). From research projects in industry, we recognize the need to consider flexible machines, flexible human workers, worker capabilities, setup and processing operations, material arrival times, complex job paths with parallel tasks for bill of material (BOM) manufacturing, sequence-dependent setup times and (partially) automated tasks in human-machine-collaboration. In recent years, there has been extensive research on metaheuristics and DRL techniques but focused on simple scheduling environments. However, there are few approaches combining metaheuristics and DRL to generate schedules more reliably and efficiently. In this paper, we first formulate a DRC-FJSSP to map complex industry requirements beyond traditional job shop models. Then we propose a scheduling framework integrating a discrete event simulation (DES) for schedule evaluation, considering parallel computing and multicriteria optimization. Here, a memetic algorithm is enriched with DRL to improve sequencing and assignment decisions. Through numerical experiments with real-world production data, we confirm that the framework generates feasible schedules efficiently and reliably for a balanced optimization of makespan (MS) and total tardiness (TT). Utilizing DRL instead of random metaheuristic operations leads to better results in fewer algorithm iterations and outperforms traditional approaches in such complex environments.
翻译:以下跨学科文章提出了一种结合深度强化学习(DRL)的模因算法,用于解决面向实际的双资源约束柔性作业车间调度问题(DRC-FJSSP)。通过工业研究项目,我们认识到需考虑柔性机器、柔性工人、工人能力、准备与加工操作、物料到达时间、面向物料清单(BOM)制造的并行任务复杂作业路径、序列依赖准备时间以及人机协作中的(部分)自动化任务。近年来,对元启发式算法和DRL技术的研究广泛,但多集中于简单调度环境。然而,将元启发式算法与DRL结合以更可靠、高效生成调度的方案仍较罕见。本文首先构建了DRC-FJSSP模型,以映射超越传统作业车间模型的复杂工业需求;随后提出一种集成离散事件仿真(DES)进行调度评估的调度框架,并考虑并行计算与多准则优化。在此框架中,利用DRL增强模因算法以改进排序与分配决策。通过基于实际生产数据的数值实验,我们证实该框架能高效可靠地生成可行调度方案,实现制造周期(MS)与总拖期时间(TT)的平衡优化。相较于随机元启发式操作,采用DRL可在更少的算法迭代中获得更优结果,并在此类复杂环境中超越传统方法。