Evaluating the performance of software for automated vehicles is predominantly driven by data collected from the real world. While professional test drivers are supported with technical means to semi-automatically annotate driving maneuvers to allow better event identification, simple data loggers in large vehicle fleets typically lack automatic and detailed event classification and hence, extra effort is needed when post-processing such data. Yet, the data quality from professional test drivers is apparently higher than the one from large fleets where labels are missing, but the non-annotated data set from large vehicle fleets is much more representative for typical, realistic driving scenarios to be handled by automated vehicles. However, while growing the data from large fleets is relatively simple, adding valuable annotations during post-processing has become increasingly expensive. In this paper, we leverage Z-order space-filling curves to systematically reduce data dimensionality while preserving domain-specific data properties, which allows us to explore even large-scale field data sets to spot interesting events orders of magnitude faster than processing time-series data directly. Furthermore, the proposed concept is based on an analytical approach, which preserves explainability for the identified events.
翻译:评估自动驾驶汽车软件性能主要依赖现实世界收集的数据。虽然专业测试驾驶员借助技术手段可半自动标注驾驶操作以提升事件识别能力,但大型车辆队列中的简易数据记录器通常缺乏自动且详细的事件分类,因此在后处理此类数据时需要额外投入。显然,专业测试驾驶员的数据质量高于缺乏标签的大型车队数据,但后者未经标注的数据集更能代表自动驾驶汽车需处理的典型真实驾驶场景。然而,尽管从大型车队扩展数据相对简单,在后处理过程中添加有价值的标注却日益昂贵。本文利用Z阶空间填充曲线系统性地降低数据维度,同时保留领域特定数据属性,这使得我们能够探索大规模现场数据集,以比直接处理时间序列数据快数个数量级的速度定位感兴趣事件。此外,所提出的概念基于分析方法,可确保已识别事件的可解释性。