Generalizable agents should adapt to diverse tasks and unseen environments beyond their training distribution. This position paper argues that such generalization requires environment scaling: expanding the distribution of executable rule-sets that agents interact with, rather than only increasing trajectories or tasks within fixed benchmarks. Current scaling practices largely focus on collecting more experience or broader task sets under fixed interaction rules, leaving agents brittle when underlying interfaces, dynamics, observations, or feedback signals change. The core challenge is therefore a world-level distribution shift: agents need systematic exposure to environments with meaningfully different executable rule-sets. To clarify this challenge, we propose a unified taxonomy that separates trajectory scaling, task scaling, and environment scaling by their primary deliverables and by what changes in the executable rule-set. Building on this taxonomy, we synthesize construction paradigms for scalable environments, contrasting programmatic generators that prioritize controllability and verifiability with generative world models that offer broader coverage and open-endedness. We further outline how environment scaling can be coupled with stateful learning mechanisms, emphasizing learned update rules for cross-environment adaptation. We conclude by discussing alternative perspectives and argue that scalable environments provide the essential substrate for measurable and controllable progress toward robust general agents.
翻译:通用型智能体应能适应多样化的任务及训练分布之外的未见环境。本文立场认为,此类泛化能力依赖于环境扩展:即扩大智能体与之交互的可执行规则集分布范围,而非仅增加固定基准中的轨迹或任务。当前扩展实践主要聚焦于在固定交互规则下收集更多经验或更广泛任务集,这导致智能体在底层接口、动力学、观测或反馈信号发生变化时变得脆弱。因此核心挑战在于世界层面的分布偏移:智能体需要系统性地暴露于具有本质差异的可执行规则集的环境中。为阐明该挑战,我们提出统一分类体系,通过主要产出物及可执行规则集的变化特征,将轨迹扩展、任务扩展与环境扩展进行区分。基于该分类体系,我们综合了可扩展环境的构建范式,对比了强调可控性与可验证性的程序化生成器与提供更广覆盖度和开放性的生成世界模型。我们进一步勾勒了环境扩展与有状态学习机制的结合路径,强调跨环境适应的学习型更新规则。最后通过讨论替代视角得出结论,论证可扩展环境为通向稳健通用智能体的可测量、可控进展提供了必要基础。