Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics. However, existing traffic-oriented multimodal benchmarks largely emphasize passive visual recognition or isolated video understanding, offering limited support for evaluating structure-aware traffic reasoning under controlled conditions. We introduce OmniTraffic, a controllable generation pipeline and benchmark for spatio-temporal traffic reasoning. Built around 12 real-world intersections reconstructed into editable 3D traffic environments and complemented by surveillance footage from two countries, OmniTraffic supports both controlled and natural-condition evaluation. It defines a three-level task hierarchy spanning scene perception, multi-view and temporal reasoning, and decision support. Using structured traffic metadata, OmniTraffic generates synchronized multi-view VQA samples covering vehicle states, lane functions, view--BEV correspondence, temporal dynamics, and signal-phase analysis, resulting in 8M VQA samples and a 3K human-verified test set. Evaluation of eleven frontier MLLMs reveals a large human--model gap, with the most pronounced failures in topology-grounded and spatio-temporal reasoning tasks. Fine-tuning a lightweight MLLM on simulated OmniTraffic data further improves performance on real-world traffic scenes, demonstrating the value of simulation-generated supervision for traffic-specific multimodal reasoning. Beyond a fixed dataset, OmniTraffic provides an extensible pipeline with configurable intersections, camera views, traffic demands, signal phases, visual conditions, and rare events.


翻译:交通场景理解要求模型具备超越目标识别的推理能力,涵盖车道拓扑、多视角几何、时态演化及信号相位语义。然而,现有面向交通的多模态基准主要侧重于被动视觉识别或孤立视频理解,在可控条件下评估结构感知型交通推理能力方面支持有限。我们提出OmniTraffic——一个面向时空交通推理的可控生成流程与基准。该基准基于12个真实交叉路口重建的可编辑3D交通环境构建,并辅以来自两个国家的监控视频数据,支持可控条件与自然条件双模式评估。其定义了涵盖场景感知、多视角与时态推理、决策支持的三级任务层级体系。通过结构化交通元数据,OmniTraffic可生成同步多视角VQA样本,覆盖车辆状态、车道功能、视角-鸟瞰图对应关系、时态动态及信号相位分析,最终产出800万条VQA样本与3000条人工校验测试集。对11个前沿多模态大语言模型的评估揭示出显著的人机性能差距,其中在拓扑 grounding 与时空推理任务中表现最为薄弱。在OmniTraffic模拟数据上微调轻量级多模态大语言模型,可进一步提升真实交通场景中的性能表现,论证了仿真监督数据在交通多模态推理中的价值。除固定数据集外,OmniTraffic还提供可扩展生成流程,支持可配置的交叉路口、相机视角、交通需求、信号相位、视觉条件及罕见事件。

0
下载
关闭预览

相关内容

【博士论文】弥合多模态基础模型与世界模型之间的鸿沟
【综述】交通流量预测,附15页论文下载
专知
23+阅读 · 2020年4月23日
出行即服务(MAAS)框架
智能交通技术
53+阅读 · 2019年5月22日
车路协同构建“通信+计算”新体系
智能交通技术
11+阅读 · 2019年3月26日
基于车路协同的群体智能协同
智能交通技术
10+阅读 · 2019年1月23日
全景分割任务介绍及其最新进展【附PPT与视频资料】
人工智能前沿讲习班
11+阅读 · 2018年12月5日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
7+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
9+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
14+阅读 · 8月7日
相关VIP内容
【博士论文】弥合多模态基础模型与世界模型之间的鸿沟
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员