Adept traffic models are critical to both planning and closed-loop simulation for autonomous vehicles (AV), and key design objectives include accuracy, diverse multimodal behaviors, interpretability, and downstream compatibility. Recently, with the advent of large language models (LLMs), an additional desirable feature for traffic models is LLM compatibility. We present Categorical Traffic Transformer (CTT), a traffic model that outputs both continuous trajectory predictions and tokenized categorical predictions (lane modes, homotopies, etc.). The most outstanding feature of CTT is its fully interpretable latent space, which enables direct supervision of the latent variable from the ground truth during training and avoids mode collapse completely. As a result, CTT can generate diverse behaviors conditioned on different latent modes with semantic meanings while beating SOTA on prediction accuracy. In addition, CTT's ability to input and output tokens enables integration with LLMs for common-sense reasoning and zero-shot generalization.
翻译:优秀的交通模型对自动驾驶车辆的规划与闭环仿真至关重要,其核心设计目标包括准确性、多样化多模态行为、可解释性及下游兼容性。近年来,随着大语言模型的出现,交通模型的另一理想特性是与LLM的兼容性。本文提出类别化交通Transformer,该模型可同时输出连续轨迹预测与标记化类别预测(车道模式、同伦类型等)。CTT最显著的特征是其完全可解释的潜在空间,这使得训练时能直接利用真实值对潜在变量进行监督,并彻底避免模式坍缩问题。因此,CTT能够基于具有语义含义的不同潜在模式生成多样化行为,同时在预测精度上超越当前最优方法。此外,CTT的输入输出标记化能力使其能集成LLM,实现常识推理与零样本泛化。