Intelligent vehicle systems require a deep understanding of the interplay between road conditions, surrounding entities, and the ego vehicle's driving behavior for safe and efficient navigation. This is particularly critical in developing countries where traffic situations are often dense and unstructured with heterogeneous road occupants. Existing datasets, predominantly geared towards structured and sparse traffic scenarios, fall short of capturing the complexity of driving in such environments. To fill this gap, we present IDD-X, a large-scale dual-view driving video dataset. With 697K bounding boxes, 9K important object tracks, and 1-12 objects per video, IDD-X offers comprehensive ego-relative annotations for multiple important road objects covering 10 categories and 19 explanation label categories. The dataset also incorporates rearview information to provide a more complete representation of the driving environment. We also introduce custom-designed deep networks aimed at multiple important object localization and per-object explanation prediction. Overall, our dataset and introduced prediction models form the foundation for studying how road conditions and surrounding entities affect driving behavior in complex traffic situations.
翻译:智能车辆系统需要深入理解道路条件、周围实体与自车驾驶行为之间的交互作用,以实现安全高效的导航。这在交通密集、非结构化且道路参与者异质性突出的发展中国家尤为关键。现有数据集主要针对结构化和稀疏的交通场景,难以捕捉此类环境中的驾驶复杂性。为填补这一空白,我们提出了IDD-X——一个大规模双视角驾驶视频数据集。该数据集包含69.7万个边界框、9000个重要目标轨迹,每个视频包含1至12个目标,为覆盖10个类别和19个解释标签类别的多个重要道路目标提供了全面的自车相对标注。数据集还整合了后视信息,以更完整地呈现驾驶环境。我们同时引入了专为多重要目标定位和逐目标解释预测而设计的定制深度网络。总体而言,本数据集及提出的预测模型为研究复杂交通场景中道路条件和周围实体如何影响驾驶行为奠定了坚实基础。