Realistic and controllable traffic simulation is a core capability that is necessary to accelerate autonomous vehicle (AV) development. However, current approaches for controlling learning-based traffic models require significant domain expertise and are difficult for practitioners to use. To remedy this, we present CTG++, a scene-level conditional diffusion model that can be guided by language instructions. Developing this requires tackling two challenges: the need for a realistic and controllable traffic model backbone, and an effective method to interface with a traffic model using language. To address these challenges, we first propose a scene-level diffusion model equipped with a spatio-temporal transformer backbone, which generates realistic and controllable traffic. We then harness a large language model (LLM) to convert a user's query into a loss function, guiding the diffusion model towards query-compliant generation. Through comprehensive evaluation, we demonstrate the effectiveness of our proposed method in generating realistic, query-compliant traffic simulations.
翻译:真实可控的交通仿真是加速自动驾驶汽车(AV)开发所需的核心能力。然而,当前控制基于学习的交通模型的方法需要大量领域专业知识,且对从业人员难以使用。为解决这一问题,我们提出CTG++——一种能够通过语言指令引导的场景级条件扩散模型。开发该模型需要应对两个挑战:一是需要具备真实性与可控性的交通模型主干,二是需要有效方法实现语言与交通模型的交互。为应对这些挑战,我们首先提出配备时空Transformer主干的场景级扩散模型,用于生成真实可控的交通流。随后利用大语言模型(LLM)将用户查询转化为损失函数,引导扩散模型生成符合查询要求的仿真结果。通过全面评估,我们证明了该方法在生成真实、符合查询要求的交通仿真方面的有效性。