Manually specifying features that capture the diversity in traffic environments is impractical. Consequently, learning-based agents cannot realize their full potential as neural motion planners for autonomous vehicles. Instead, this work proposes to learn which features are task-relevant. Given its immediate relevance to motion planning, our proposed architecture encodes the probabilistic occupancy map as a proxy for obtaining pre-trained state representations. By leveraging a map-aware graph formulation of the environment, our agent-centric encoder generalizes to arbitrary road networks and traffic situations. We show that our approach significantly improves the downstream performance of a reinforcement learning agent operating in urban traffic environments.
翻译:人工指定能捕捉交通环境多样性的特征是不切实际的。因此,基于学习的智能体无法充分发挥其作为自动驾驶神经运动规划器的潜力。相反,本工作提出学习哪些特征与任务相关。鉴于其与运动规划的直接关联,我们提出的架构将概率占位图编码为获取预训练状态表征的代理指标。通过利用对环境的地图感知图结构建模,以智能体为中心的编码器可泛化至任意道路网络和交通场景。实验表明,本方法显著提升了在城区交通环境中运行的强化学习智能体的下游性能。