Developing a general algorithm that learns to solve tasks across a wide range of applications has been a fundamental challenge in artificial intelligence. Although current reinforcement learning algorithms can be readily applied to tasks similar to what they have been developed for, configuring them for new application domains requires significant human expertise and experimentation. We present DreamerV3, a general algorithm that outperforms specialized methods across over 150 diverse tasks, with a single configuration. Dreamer learns a model of the environment and improves its behavior by imagining future scenarios. Robustness techniques based on normalization, balancing, and transformations enable stable learning across domains. Applied out of the box, Dreamer is the first algorithm to collect diamonds in Minecraft from scratch without human data or curricula. This achievement has been posed as a significant challenge in artificial intelligence that requires exploring farsighted strategies from pixels and sparse rewards in an open world. Our work allows solving challenging control problems without extensive experimentation, making reinforcement learning broadly applicable.
翻译:开发一种能够学习解决广泛领域任务的通用算法一直是人工智能领域的基本挑战。尽管当前的强化学习算法可以轻松应用于与其开发任务相似的任务,但为新应用领域配置这些算法需要大量人类专业知识和实验。我们提出了DreamerV3,这是一种通用算法,在超过150个不同任务中以单一配置超越了专门方法。Dreamer通过想象未来情景来学习环境模型并改进其行为。基于归一化、平衡和变换的鲁棒性技术使得跨领域稳定学习成为可能。Dreamer是首个无需人类数据或课程即可从零开始收集《我的世界》中钻石的算法——这一成就被视为人工智能领域的重大挑战,要求从像素和稀疏奖励中探索具有前瞻性的策略。我们的工作使解决具有挑战性的控制问题无需大量实验,从而让强化学习具有广泛适用性。