One of the grand challenges of reinforcement learning is the ability to generalize to new tasks. However, general agents require a set of rich, diverse tasks to train on. Designing a `foundation environment' for such tasks is tricky -- the ideal environment would support a range of emergent phenomena, an expressive task space, and fast runtime. To take a step towards addressing this research bottleneck, this work presents Powderworld, a lightweight yet expressive simulation environment running directly on the GPU. Within Powderworld, two motivating challenges distributions are presented, one for world-modelling and one for reinforcement learning. Each contains hand-designed test tasks to examine generalization. Experiments indicate that increasing the environment's complexity improves generalization for world models and certain reinforcement learning agents, yet may inhibit learning in high-variance environments. Powderworld aims to support the study of generalization by providing a source of diverse tasks arising from the same core rules.
翻译:强化学习的重大挑战之一在于对新任务的泛化能力。然而,通用智能体需要基于丰富多样的任务分布进行训练。为此类任务设计"基础环境"颇具难度——理想环境应支持多种涌现现象、富有表现力的任务空间以及快速运行时。为突破这一研究瓶颈,本文提出"粉世界"——一个直接运行在GPU上的轻量级且富有表现力的模拟环境。基于该环境,我们构建了两种挑战性任务分布:一种面向世界建模,另一种面向强化学习。每种分布均包含人工设计的测试任务以检验泛化能力。实验表明,提升环境复杂度能改善世界模型与特定强化学习智能体的泛化性能,但在高方差环境中可能抑制学习效果。粉世界旨在通过提供由同一核心规则衍生的多样化任务分布,支持泛化研究。