The reward system is one of the fundamental drivers of animal behaviors and is critical for survival and reproduction. Despite its importance, the problem of how the reward system has evolved is underexplored. In this paper, we try to replicate the evolution of biologically plausible reward functions and investigate how environmental conditions affect evolved rewards' shape. For this purpose, we developed a population-based decentralized evolutionary simulation framework, where agents maintain their energy level to live longer and produce more children. Each agent inherits its reward function from its parent subject to mutation and learns to get rewards via reinforcement learning throughout its lifetime. Our results show that biologically reasonable positive rewards for food acquisition and negative rewards for motor action can evolve from randomly initialized ones. However, we also find that the rewards for motor action diverge into two modes: largely positive and slightly negative. The emergence of positive motor action rewards is surprising because it can make agents too active and inefficient in foraging. In environments with poor and poisonous foods, the evolution of rewards for less important foods tends to be unstable, while rewards for normal foods are still stable. These results demonstrate the usefulness of our simulation environment and energy-dependent birth and death model for further studies of the origin of reward systems.
翻译:奖励系统是驱动动物行为的基本动力之一,对生存与繁殖至关重要。尽管其重要性不言而喻,但关于奖励系统如何演化的问题仍未得到充分探索。本文尝试复现具有生物学合理性的奖励函数的演化过程,并研究环境条件如何影响演化奖励的形态。为此,我们开发了一个基于种群的分散式演化模拟框架,其中智能体通过维持自身能量水平以延长寿命并繁衍更多后代。每个智能体通过变异继承其父代的奖励函数,并在其生命周期中通过强化学习学会获取奖励。研究结果表明,对于获取食物的正向奖励与对于运动行为的负向奖励能够从随机初始化的奖励函数中演化产生,这与生物学观察相符。然而,我们也发现运动行为奖励会分化为两种模式:显著正向与轻微负向。正向运动行为奖励的出现令人意外,因为它可能导致智能体过度活跃从而降低觅食效率。在食物匮乏且存在有毒食物的环境中,针对次要食物的奖励演化往往不稳定,而对正常食物的奖励仍保持稳定。这些结果证明了我们的模拟环境及能量依赖的生死模型在研究奖励系统起源问题上的有效性。