The foundational role of datasets in defining the capabilities of deep learning models has led to their rapid proliferation. At the same time, published research focusing on the process of dataset development for environment perception in automated driving has been scarce, thereby reducing the applicability of openly available datasets and impeding the development of effective environment perception systems. Sensor-based, mapless automated driving is one of the contexts where this limitation is evident. While leveraging real-time sensor data, instead of pre-defined HD maps promises enhanced adaptability and safety by effectively navigating unexpected environmental changes, it also increases the demands on the scope and complexity of the information provided by the perception system. To address these challenges, we propose a scenario- and capability-based approach to dataset development. Grounded in the principles of ISO 21448 (safety of the intended functionality, SOTIF), extended by ISO/TR 4804, our approach facilitates the structured derivation of dataset requirements. This not only aids in the development of meaningful new datasets but also enables the effective comparison of existing ones. Applying this methodology to a broad range of existing lane detection datasets, we identify significant limitations in current datasets, particularly in terms of real-world applicability, a lack of labeling of critical features, and an absence of comprehensive information for complex driving maneuvers.
翻译:数据集在定义深度学习模型能力中的基础性作用,推动了其快速涌现。然而,针对自动驾驶环境感知数据集开发流程的公开研究仍显匮乏,这既降低了公开数据集的适用性,也阻碍了高效环境感知系统的研发。基于传感器、无地图的自动驾驶正是这一局限性的典型场景:尽管利用实时传感器数据而非预定义高精地图,有望通过有效应对突发环境变化提升适应性与安全性,但同时也对感知系统所提供信息的广度与复杂性提出了更高要求。为应对这些挑战,我们提出了一种基于场景与能力的数据集开发方法。该方法以ISO 21448(预期功能安全,SOTIF)为理论基础,并扩展引入ISO/TR 4804相关原则,能结构化地推导数据集需求。这不仅有助于开发具有实际意义的新数据集,还可实现现有数据集的有效对比。通过将该方法论应用于现有车道检测数据集进行广泛评估,我们发现当前数据集存在显著局限性,具体表现为:真实世界适用性不足、关键特征标注缺失,以及复杂驾驶场景中综合信息的缺乏。