Predicting temporally consistent road users' trajectories in a multi-agent setting is a challenging task due to unknown characteristics of agents and their varying intentions. Besides using semantic map information and modeling interactions, it is important to build an effective mechanism capable of reasoning about behaviors at different levels of granularity. To this end, we propose Dynamic goal quErieS with temporal Transductive alIgNmEnt (DESTINE) method. Unlike past arts, our approach 1) dynamically predicts agents' goals irrespective of particular road structures, such as lanes, allowing the method to produce a more accurate estimation of destinations; 2) achieves map compliant predictions by generating future trajectories in a coarse-to-fine fashion, where the coarser predictions at a lower frame rate serve as intermediate goals; and 3) uses an attention module designed to temporally align predicted trajectories via masked attention. Using the common Argoverse benchmark dataset, we show that our method achieves state-of-the-art performance on various metrics, and further investigate the contributions of proposed modules via comprehensive ablation studies.
翻译:在多智能体场景中预测时间一致的道路使用者轨迹是一项具有挑战性的任务,原因在于智能体未知的特性及其动态变化的意图。除利用语义地图信息和建模交互外,构建能够对不同粒度行为进行推理的有效机制至关重要。为此,我们提出基于时间传导对齐的动态目标查询(DESTINE)方法。与现有方法不同,本方法:1)动态预测智能体目标而不依赖特定道路结构(如车道),从而更准确地估计目的地;2)通过从粗到细的方式生成未来轨迹实现地图合规预测,其中低帧率下的粗粒度预测作为中间目标;3)采用注意力模块,通过掩码注意力机制对预测轨迹进行时间对齐。基于通用Argoverse基准数据集,我们证明了该方法在多项指标上达到当前最优性能,并通过全面的消融实验进一步探究了所提模块的贡献。