Extracting actionable insights from long-duration urban videos is often labor-intensive: analysts must manually sift through raw footage to pinpoint target events or uncover broader behavioral trends. In this work, we present URBANCLIPATLAS, a visual analytics system for exploring long urban videos recorded at street intersections. URBANCLIPATLAS combines retrieval-augmented generation (RAG), taxonomy-aware entity extraction, and video grounding to support event retrieval and interpretation. The system segments extended recordings into short clips, generates textual descriptions with a vision-language model, and indexes them for semantic retrieval. A knowledge graph maps entities and relations from LLM answers onto a domain-specific taxonomy and aligns them with detected objects and trajectories to support visual grounding and verification. URBANCLIPATLAS supports scene retrieval through an augmented chat-based interface and improves scene interpretation by tightly aligning textual outputs with video evidence. This design strengthens the connection between textual reasoning and visual evidence, reducing the effort required to validate model outputs and refine hypotheses. We demonstrate the usefulness of URBANCLIPATLAS on the StreetAware dataset through two case studies involving hazardous scenarios and crossing dynamics at street intersections. URBANCLIPATLAS helps analysts reason about safety- and mobility-related patterns across large urban video collections.


翻译:从长时间城市视频中提取可操作洞察通常需要大量人工劳动:分析人员必须手动筛选原始录像以定位目标事件或发现更广泛的行为趋势。本文提出URBANCLIPATLAS——一个用于探索街道交叉口记录的长时间城市视频的可视分析系统。URBANCLIPATLAS结合检索增强生成、分类感知实体提取与视频定位技术,支持事件检索与解读。该系统将长时录视频分段为短片段,利用视觉-语言模型生成文本描述并建立索引以实现语义检索。知识图谱将大语言模型回答中的实体与关系映射到领域特定分类体系,并将其与检测对象及轨迹对齐以支持视觉定位与验证。URBANCLIPATLAS通过增强型对话界面支持场景检索,并通过将文本输出与视频证据紧密对齐来提升场景解读效果。该设计强化了文本推理与视觉证据之间的联系,减少了验证模型输出与完善假设所需的工作量。我们通过两个涉及十字路口危险场景与穿越动态的案例研究,在StreetAware数据集上展示了URBANCLIPATLAS的实用性。该系统帮助分析人员对大规模城市视频集合中与安全及出行相关的模式进行推理。

0
下载
关闭预览

相关内容

【博士论文】面向城市环境的可解释计算机视觉
专知会员服务
17+阅读 · 4月18日
城市大数据认知计算研究与应用进展
专知会员服务
29+阅读 · 2024年7月18日
【CVPR2024】OmniViD: 一个用于通用视频理解的生成框架
专知会员服务
25+阅读 · 2024年3月27日
流行病数据可视分析综述
专知会员服务
27+阅读 · 2022年3月21日
专知会员服务
53+阅读 · 2020年12月19日
艾瑞咨询2019中国智慧城市发展报告,附PPT下载
智能交通技术
25+阅读 · 2019年4月18日
一种轻量级在线多目标车辆跟踪方法
极市平台
15+阅读 · 2018年8月18日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
《多域冲突比较支持模型》60页
专知会员服务
8+阅读 · 8月7日
面向2027年及未来的海军情报改革
专知会员服务
5+阅读 · 8月5日
相关VIP内容
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员