The success of Deep Reinforcement Learning (DRL) is largely attributed to utilizing Artificial Neural Networks (ANNs) as function approximators. Recent advances in neuroscience have unveiled that the human brain achieves efficient reward-based learning, at least by integrating spiking neurons with spatial-temporal dynamics and network topologies with biologically-plausible connectivity patterns. This integration process allows spiking neurons to efficiently combine information across and within layers via nonlinear dendritic trees and lateral interactions. The fusion of these two topologies enhances the network's information-processing ability, crucial for grasping intricate perceptions and guiding decision-making procedures. However, ANNs and brain networks differ significantly. ANNs lack intricate dynamical neurons and only feature inter-layer connections, typically achieved by direct linear summation, without intra-layer connections. This limitation leads to constrained network expressivity. To address this, we propose a novel alternative for function approximator, the Biologically-Plausible Topology improved Spiking Actor Network (BPT-SAN), tailored for efficient decision-making in DRL. The BPT-SAN incorporates spiking neurons with intricate spatial-temporal dynamics and introduces intra-layer connections, enhancing spatial-temporal state representation and facilitating more precise biological simulations. Diverging from the conventional direct linear weighted sum, the BPT-SAN models the local nonlinearities of dendritic trees within the inter-layer connections. For the intra-layer connections, the BPT-SAN introduces lateral interactions between adjacent neurons, integrating them into the membrane potential formula to ensure accurate spike firing.
翻译:深度强化学习(DRL)的成功主要归功于利用人工神经网络(ANNs)作为函数逼近器。神经科学的最新进展揭示,人脑至少通过整合具有时空动态特性的脉冲神经元和具有生物合理连接模式的网络拓扑结构,实现了高效的基于奖励的学习。这一整合过程使得脉冲神经元能够通过非线性树突树和侧向相互作用,高效地跨层和层内组合信息。这两种拓扑结构的融合增强了网络的信息处理能力,对于理解复杂感知和指导决策过程至关重要。然而,人工神经网络与大脑网络存在显著差异。人工神经网络缺乏复杂的动态神经元,仅具备层间连接,通常通过直接线性求和实现,而缺乏层内连接。这一局限性导致网络表达能力受限。为解决此问题,我们提出了一种新型函数逼近器替代方案——生物启发的拓扑优化脉冲演员网络(BPT-SAN),专门用于DRL中的高效决策。BPT-SAN包含具有复杂时空动态特性的脉冲神经元,并引入层内连接,增强了时空状态表示能力,并实现更精确的生物模拟。与传统直接线性加权求和不同,BPT-SAN在层间连接中建模了树突树的局部非线性特性。对于层内连接,BPT-SAN引入了相邻神经元之间的侧向相互作用,并将其整合到膜电位公式中,以确保精准的脉冲发放。