Bird's eye view (BEV) perception is becoming increasingly important in the field of autonomous driving. It uses multi-view camera data to learn a transformer model that directly projects the perception of the road environment onto the BEV perspective. However, training a transformer model often requires a large amount of data, and as camera data for road traffic are often private, they are typically not shared. Federated learning offers a solution that enables clients to collaborate and train models without exchanging data but model parameters. In this paper, we introduce FedBEVT, a federated transformer learning approach for BEV perception. In order to address two common data heterogeneity issues in FedBEVT: (i) diverse sensor poses, and (ii) varying sensor numbers in perception systems, we propose two approaches -- Federated Learning with Camera-Attentive Personalization (FedCaP) and Adaptive Multi-Camera Masking (AMCM), respectively. To evaluate our method in real-world settings, we create a dataset consisting of four typical federated use cases. Our findings suggest that FedBEVT outperforms the baseline approaches in all four use cases, demonstrating the potential of our approach for improving BEV perception in autonomous driving.
翻译:鸟瞰视角(BEV)感知在自动驾驶领域正变得越来越重要。它利用多视角摄像头数据学习一个Transformer模型,直接将道路环境的感知投影到BEV视角上。然而,训练Transformer模型通常需要大量数据,而道路交通的摄像头数据往往涉及隐私,因此通常不共享。联邦学习提供了一种解决方案,使得客户端能够在不交换数据而仅交换模型参数的情况下协作训练模型。本文提出了FedBEVT——一种面向BEV感知的联邦Transformer学习方法。为解决FedBEVT中两类常见的数据异构性问题:(i)传感器姿态多样性及(ii)感知系统中传感器数量的变化,我们分别提出了两种方法——基于摄像头注意力个性化的联邦学习(FedCaP)和自适应多摄像头掩码(AMCM)。为在真实场景中评估我们的方法,我们构建了一个包含四种典型联邦用例的数据集。实验结果表明,FedBEVT在所有四种用例中均优于基线方法,展示了该方法在提升自动驾驶中BEV感知性能方面的潜力。