In autonomous driving tasks, scene understanding is the first step towards predicting the future behavior of the surrounding traffic participants. Yet, how to represent a given scene and extract its features are still open research questions. In this study, we propose a novel text-based representation of traffic scenes and process it with a pre-trained language encoder. First, we show that text-based representations, combined with classical rasterized image representations, lead to descriptive scene embeddings. Second, we benchmark our predictions on the nuScenes dataset and show significant improvements compared to baselines. Third, we show in an ablation study that a joint encoder of text and rasterized images outperforms the individual encoders confirming that both representations have their complementary strengths.
翻译:在自动驾驶任务中,场景理解是预测周围交通参与者未来行为的第一步。然而,如何表征给定场景并提取其特征仍是尚待研究的问题。本研究提出一种基于文本的新型交通场景表征方法,并利用预训练语言编码器对其进行处理。首先,我们证明了基于文本的表征与经典的栅格图像表征相结合,能够生成描述性场景嵌入。其次,我们在nuScenes数据集上对预测结果进行基准测试,结果表明相比基线方法有显著改进。第三,消融实验表明,文本与栅格图像的联合编码器优于单独编码器,证实两种表征具有互补优势。