Disruption recovery in industrial assembly lines requires timely decisions under machine faults, worker absence, and emergency orders. Existing methods either rely on rigid handcrafted recovery logic or learn adaptive policies that do not readily exploit heterogeneous external recovery knowledge at decision time to reduce abnormal recovery time (ART) and preserve on-time delivery (OTD). To address this gap, we propose a phase-aware guidance injection framework that augments a trained recurrent MAPPO (RMAPPO) scheduling policy through logit-level action bias during evaluation. The framework provides a unified decision-time interface for rule-based, replay-based, and online LLM-based guidance, while activating intervention only during abnormal and recovery phases. Experiments on a custom AssemblyLineEnv show that high-quality rule guidance yields the strongest gains, replay-based guidance degrades smoothly under imperfect availability, and online LLM guidance still provides useful intermediate improvements. These results show that decision-time guidance injection can exploit heterogeneous recovery hints without redesigning the actor.
翻译:工业装配线的中断恢复要求在面对机器故障、人员缺勤和紧急订单时做出及时决策。现有方法要么依赖僵化的人工设计恢复逻辑,要么学习自适应策略,但这些策略在决策时无法充分利用异构的外部恢复知识来减少异常恢复时间(ART)并保持按时交付(OTD)。为弥补这一不足,我们提出了一种相位感知引导注入框架,该框架通过在评估期间施加logit级动作偏置,增强已训练的循环MAPPO(RMAPPO)调度策略。该框架为基于规则、基于重放和基于在线LLM的引导提供了统一的决策时接口,同时仅在异常相位和恢复相位激活干预。在自建AssemblyLineEnv环境上的实验表明,高质量规则引导带来了最强的性能提升,基于重放的引导在不完美可用性下性能平滑下降,而在线LLM引导仍能提供有用的中间改进。这些结果表明,决策时引导注入无需重新设计智能体即可利用异构恢复提示。