Trajectory prediction in autonomous driving relies on accurate representation of all relevant contexts of the driving scene including traffic participants, road topology, traffic signs as well as their semantic relations to each other. Despite increased attention to this issue, most approaches in trajectory prediction do not consider all of these factors sufficiently. This paper describes a method SemanticFormer to predict multimodal trajectories by reasoning over a semantic traffic scene graph using a hybrid approach. We extract high-level information in the form of semantic meta-paths from a knowledge graph which is then processed by a novel pipeline based on multiple attention mechanisms to predict accurate trajectories. The proposed architecture comprises a hierarchical heterogeneous graph encoder, which can capture spatio-temporal and relational information across agents and between agents and road elements, and a predictor that fuses the different encodings and decodes trajectories with probabilities. Finally, a refinement module evaluates permitted meta-paths of trajectories and speed profiles to obtain final predicted trajectories. Evaluation of the nuScenes benchmark demonstrates improved performance compared to the state-of-the-art methods.
翻译:摘要:自动驾驶中的轨迹预测依赖于对驾驶场景所有相关语境要素的精确表示,包括交通参与者、道路拓扑、交通标志及其之间的语义关联。尽管该问题日益受到关注,但现有大多数轨迹预测方法未能充分兼顾所有这些因素。本文提出SemanticFormer方法,通过混合推理机制对语义交通场景图进行推理,实现多模态轨迹预测。我们从知识图谱中提取语义元路径形式的高层信息,随后通过基于多重注意力机制的新型流水线进行处理,以生成精确轨迹。所提出的架构包含层级式异构图编码器(可捕捉智能体间及智能体与道路要素间的时空与关联信息)和预测器(融合不同编码结果并解码轨迹概率)。最终由精炼模块评估轨迹的可行元路径与速度剖面,获得最终预测轨迹。在nuScenes基准上的评估表明,该方法相较现有最优方法获得了性能提升。