Flexible manufacturing has given rise to complex scheduling problems such as the flexible job shop scheduling problem (FJSP). In FJSP, operations can be processed on multiple machines, leading to intricate relationships between operations and machines. Recent works have employed deep reinforcement learning (DRL) to learn priority dispatching rules (PDRs) for solving FJSP. However, the quality of solutions still has room for improvement relative to that by the exact methods such as OR-Tools. To address this issue, this paper presents a novel end-to-end learning framework that weds the merits of self-attention models for deep feature extraction and DRL for scalable decision-making. The complex relationships between operations and machines are represented precisely and concisely, for which a dual-attention network (DAN) comprising several interconnected operation message attention blocks and machine message attention blocks is proposed. The DAN exploits the complicated relationships to construct production-adaptive operation and machine features to support high-quality decisionmaking. Experimental results using synthetic data as well as public benchmarks corroborate that the proposed approach outperforms both traditional PDRs and the state-of-the-art DRL method. Moreover, it achieves results comparable to exact methods in certain cases and demonstrates favorable generalization ability to large-scale and real-world unseen FJSP tasks.
翻译:柔性制造引发了诸如柔性作业车间调度问题(FJSP)等复杂调度问题。在FJSP中,工序可在多台机器上加工,导致工序与机器之间存在错综复杂的关联。近期研究采用深度强化学习(DRL)学习优先调度规则(PDRs)以求解FJSP,但相较于OR-Tools等精确方法,其解质量仍有提升空间。为此,本文提出一种新颖的端到端学习框架,融合自注意力模型在深层特征提取方面的优势与DRL在可扩展决策方面的特性。通过精确简明地表征工序与机器间的复杂关系,提出由多个互联的工序消息注意力块与机器消息注意力块构成的双重注意力网络(DAN)。DAN利用这些复杂关系构建生产自适应的工序与机器特征,以支持高质量决策。基于合成数据与公开基准的实验结果表明,所提方法在性能上优于传统PDRs及最先进的DRL方法,并在特定案例中达到与精确方法可比的效果,展现出对大规模及真实未知FJSP任务的良好泛化能力。