The accurate modeling of dynamics in interactive environments is critical for successful long-range prediction. Such a capability could advance Reinforcement Learning (RL) and Planning algorithms, but achieving it is challenging. Inaccuracies in model estimates can compound, resulting in increased errors over long horizons. We approach this problem from the lens of Koopman theory, where the nonlinear dynamics of the environment can be linearized in a high-dimensional latent space. This allows us to efficiently parallelize the sequential problem of long-range prediction using convolution, while accounting for the agent's action at every time step. Our approach also enables stability analysis and better control over gradients through time. Taken together, these advantages result in significant improvement over the existing approaches, both in the efficiency and the accuracy of modeling dynamics over extended horizons. We also report promising experimental results in dynamics modeling for the scenarios of both model-based planning and model-free RL.
翻译:在交互环境中准确建模动力学对于实现成功的长期预测至关重要。这种能力能够推进强化学习和规划算法的发展,但实现这一目标极具挑战性。模型估计中的误差会累积叠加,导致长期预测中误差持续增大。我们从库普曼理论的视角处理该问题——在该理论框架下,环境的非线性动力学可以在高维潜在空间中被线性化。这使得我们能够利用卷积高效并行处理长期预测的序列问题,同时考虑智能体在每个时间步的动作。我们的方法还支持稳定性分析,并能更好地控制时间维度上的梯度传播。综合这些优势,使得我们的方法在扩展时间维度的动力学建模效率和准确性上均显著优于现有方法。我们在基于模型的规划和无模型强化学习这两种场景下的动力学建模实验中,也报告了令人鼓舞的成果。