This paper introduces a novel approach, Decision Theory-guided Deep Reinforcement Learning (DT-guided DRL), to address the inherent cold start problem in DRL. By integrating decision theory principles, DT-guided DRL enhances agents' initial performance and robustness in complex environments, enabling more efficient and reliable convergence during learning. Our investigation encompasses two primary problem contexts: the cart pole and maze navigation challenges. Experimental results demonstrate that the integration of decision theory not only facilitates effective initial guidance for DRL agents but also promotes a more structured and informed exploration strategy, particularly in environments characterized by large and intricate state spaces. The results of experiment demonstrate that DT-guided DRL can provide significantly higher rewards compared to regular DRL. Specifically, during the initial phase of training, the DT-guided DRL yields up to an 184% increase in accumulated reward. Moreover, even after reaching convergence, it maintains a superior performance, ending with up to 53% more reward than standard DRL in large maze problems. DT-guided DRL represents an advancement in mitigating a fundamental challenge of DRL by leveraging functions informed by human (designer) knowledge, setting a foundation for further research in this promising interdisciplinary domain.
翻译:本文提出了一种新颖的方法——决策理论引导的深度强化学习(DT-guided DRL),以解决深度强化学习(DRL)中固有的冷启动问题。通过整合决策理论原理,DT-guided DRL提升了智能体在复杂环境中的初始性能和鲁棒性,使其在学习过程中能够实现更高效、更可靠的收敛。我们的研究涵盖了两个主要问题场景:倒立摆和迷宫导航挑战。实验结果表明,决策理论的整合不仅为DRL智能体提供了有效的初始引导,还促进了更具结构性和信息性的探索策略,尤其是在状态空间庞大且复杂的环境中。实验结果显示,与常规DRL相比,DT-guided DRL能够获得显著更高的奖励。具体而言,在训练的初始阶段,DT-guided DRL的累积奖励提升了高达184%。此外,即使在达到收敛后,它在大型迷宫问题中的奖励仍比标准DRL高出多达53%。DT-guided DRL通过利用基于人类(设计者)知识的函数,在缓解DRL的一个基本挑战方面取得了进展,为这一有前景的跨学科领域的进一步研究奠定了基础。