Fleets of networked autonomous vehicles (AVs) collect terabytes of sensory data, which is often transmitted to central servers (the ''cloud'') for training machine learning (ML) models. Ideally, these fleets should upload all their data, especially from rare operating contexts, in order to train robust ML models. However, this is infeasible due to prohibitive network bandwidth and data labeling costs. Instead, we propose a cooperative data sampling strategy where geo-distributed AVs collaborate to collect a diverse ML training dataset in the cloud. Since the AVs have a shared objective but minimal information about each other's local data distribution and perception model, we can naturally cast cooperative data collection as an $N$-player mathematical game. We show that our cooperative sampling strategy uses minimal information to converge to a centralized oracle policy with complete information about all AVs. Moreover, we theoretically characterize the performance benefits of our game-theoretic strategy compared to greedy sampling. Finally, we experimentally demonstrate that our method outperforms standard benchmarks by up to $21.9\%$ on 4 perception datasets, including for autonomous driving in adverse weather conditions. Crucially, our experimental results on real-world datasets closely align with our theoretical guarantees.
翻译:联网自主车辆(AV)车队会收集海量传感数据,这些数据通常被传输至中央服务器(即“云端”)以训练机器学习(ML)模型。理想情况下,这些车队应上传全部数据(尤其是来自罕见运行场景的数据),从而训练出鲁棒的机器学习模型。然而,由于网络带宽和数据标注成本过高,这一方案并不可行。为此,我们提出一种协同数据采样策略,使地理分布的自主车辆协作在云端收集多样化的机器学习训练数据集。由于自主车辆具有共享目标,但对彼此本地数据分布和感知模型的了解极少,我们自然地将协同数据收集建模为$N$玩家数学博弈。我们证明,该协同采样策略仅需极少量信息即可收敛至具备所有自主车辆完整信息的集中式最优策略。此外,我们从理论上刻画了与贪心采样相比,我们博弈论策略的性能优势。最后,实验表明,我们的方法在4个感知数据集(包括恶劣天气条件下的自动驾驶场景)上的表现优于标准基准方法多达$21.9\%$。关键的是,我们在真实数据集上的实验结果与理论保障高度吻合。