Inspired by recent progress in multi-agent Reinforcement Learning (RL), in this work we examine the collective intelligent behaviour of theoretical universal agents by introducing a weighted mixture operation. Given a weighted set of agents, their weighted mixture is a new agent whose expected total reward in any environment is the corresponding weighted average of the original agents' expected total rewards in that environment. Thus, if RL agent intelligence is quantified in terms of performance across environments, the weighted mixture's intelligence is the weighted average of the original agents' intelligences. This operation enables various interesting new theorems that shed light on the geometry of RL agent intelligence, namely: results about symmetries, convex agent-sets, and local extrema. We also show that any RL agent intelligence measure based on average performance across environments, subject to certain weak technical conditions, is identical (up to a constant factor) to performance within a single environment dependent on said intelligence measure.
翻译:受多智能体强化学习(RL)最新进展的启发,本文通过引入加权混合操作,研究了理论通用智能体的集体智能行为。给定一组加权智能体,其加权混合是一个新智能体,该智能体在任何环境中的期望总回报等于原始智能体在该环境中期望总回报的相应加权平均值。因此,若RL智能体的智能以跨环境性能来衡量,则加权混合智能体的智能即为原始智能体智能的加权平均值。该操作催生了一系列有趣的新定理,揭示了RL智能体智能的几何特性,具体包括:对称性、凸智能体集以及局部极值相关结论。我们还证明,基于跨环境平均性能的任何RL智能体智能度量(在满足一定弱技术条件的前提下)与依赖于该度量的单一环境中的性能(相差一个常数因子)是等价的。