Learned driving agents often degrade when deployed in unseen environments. This paper studies a deliberately bounded instance of that problem in the CARLA simulator: zero-shot transfer of a closed-loop fixed-route driving agent from Town05 and Town06 to unseen Town03 and Town04. The study isolates structural town shift by keeping weather fixed to ClearNoon and removing traffic and pedestrians. We build on a Dreamer-style latent world-model agent and add two training-only auxiliary losses: multi-horizon prediction of future visual-semantic embeddings along imagined rollouts and town-adversarial supervision on a semantic projection of the recurrent latent state. A causal context feature conditions the semantic rollout predictor, while the actor and critic retain the standard control feature. The policy receives no navigation command, route polyline, goal pose, or map input; the reference route is used only by the environment for reward, progress, success, and termination. Across the evaluated held-out towns, the proposed model achieves the highest mean success rate among the included Dreamer-family methods. Secondary safety and lane-keeping metrics are mixed across towns. These results support a bounded conclusion: in this controlled fixed-weather CARLA setting, semantic rollout supervision combined with town-adversarial regularization improves mean held-out-town route completion.
翻译:学习型驾驶智能体在部署到未见过的环境时往往会性能下降。本文在 CARLA 模拟器中研究了该问题的一个刻意限定的实例:将闭环固定路线驾驶智能体从 Town05 和 Town06 零样本迁移到未见过的 Town03 和 Town04。研究通过保持天气固定为晴朗正午并移除交通和行人,隔离了结构性城镇偏移。我们基于 Dreamer 风格的潜在世界模型智能体,并添加两个仅训练时使用的辅助损失:沿想象展开的未来视觉语义嵌入的多时域预测,以及基于循环潜状态语义投影的城镇对抗监督。因果上下文特征对语义展开预测器进行条件约束,而演员与评论家网络保留标准控制特征。策略不接收任何导航指令、路线折线、目标位姿或地图输入;参考路线仅由环境用于奖励、进度、成功和终止判定。在评估的所有未见城镇中,所提模型在包含的 Dreamer 系列方法中实现了最高的平均成功率。次要安全性和车道保持指标在各城镇表现不一。这些结果支持一个有限结论:在此受控固定天气的 CARLA 设定下,语义展开监督与城镇对抗正则化相结合提升了未见城镇的路线完成平均成功率。