The limited priors required by neural networks make them the dominating choice to encode and learn policies using reinforcement learning (RL). However, they are also black-boxes, making it hard to understand the agent's behaviour, especially when working on the image level. Therefore, neuro-symbolic RL aims at creating policies that are interpretable in the first place. Unfortunately, interpretability is not explainability. To achieve both, we introduce Neurally gUided Differentiable loGic policiEs (NUDGE). NUDGE exploits trained neural network-based agents to guide the search of candidate-weighted logic rules, then uses differentiable logic to train the logic agents. Our experimental evaluation demonstrates that NUDGE agents can induce interpretable and explainable policies while outperforming purely neural ones and showing good flexibility to environments of different initial states and problem sizes.
翻译:神经网络所需的先验知识较少,使其成为通过强化学习(RL)编码和学习策略的主要选择。然而,神经网络也是黑箱模型,难以理解智能体的行为,尤其是在图像层面工作时。因此,神经符号RL旨在创建一开始就可解释的策略。但不幸的是,可解释性并不等同于可说明性。为实现两者兼得,我们提出了神经引导可微逻辑策略(NUDGE)。NUDGE利用基于神经网络的训练智能体来引导候选加权逻辑规则的搜索,然后通过可微逻辑训练逻辑智能体。我们的实验评估表明,NUDGE智能体能够生成可解释且可说明的策略,同时性能优于纯神经网络智能体,并在不同初始状态和问题规模的环境中展现出良好的灵活性。