Modern AI applications involving video, such as video-text alignment, video search, and video captioning, benefit from a fine-grained understanding of video semantics. Existing approaches for video understanding are either data-hungry and need low-level annotation, or are based on general embeddings that are uninterpretable and can miss important details. We propose LASER, a neuro-symbolic approach that learns semantic video representations by leveraging logic specifications that can capture rich spatial and temporal properties in video data. In particular, we formulate the problem in terms of alignment between raw videos and specifications. The alignment process efficiently trains low-level perception models to extract a fine-grained video representation that conforms to the desired high-level specification. Our pipeline can be trained end-to-end and can incorporate contrastive and semantic loss functions derived from specifications. We evaluate our method on two datasets with rich spatial and temporal specifications: 20BN-Something-Something and MUGEN. We demonstrate that our method not only learns fine-grained video semantics but also outperforms existing baselines on downstream tasks such as video retrieval.


翻译:摘要:现代涉及视频的人工智能应用(如视频-文本对齐、视频搜索和视频字幕生成)受益于对视频语义的细粒度理解。现有的视频理解方法要么数据需求量大且需要低级标注,要么基于不可解释且可能遗漏重要细节的通用嵌入。我们提出LASER,一种通过利用能够捕获视频数据中丰富时空属性的逻辑规范来学习语义视频表示的神经符号方法。具体而言,我们将问题表述为原始视频与规范之间的对齐。该对齐过程高效训练低级感知模型,以提取符合所需高级规范的细粒度视频表示。我们的流水线可进行端到端训练,并能整合从规范中导出的对比损失与语义损失函数。我们在两个具有丰富时空规范的数据集(20BN-Something-Something和MUGEN)上评估了该方法。结果表明,我们的方法不仅学习了细粒度视频语义,而且在视频检索等下游任务中优于现有基线方法。

0
下载
关闭预览

相关内容

【AAAI2022】(2.5+1)D时空场景图用于视频问答
专知会员服务
24+阅读 · 2022年2月21日
专知会员服务
16+阅读 · 2021年10月4日
【ACL2020放榜!】事件抽取、关系抽取、NER、Few-Shot 相关论文整理
深度学习自然语言处理
18+阅读 · 2020年5月22日
Transferring Knowledge across Learning Processes
CreateAMind
29+阅读 · 2019年5月18日
逆强化学习-学习人先验的动机
CreateAMind
16+阅读 · 2019年1月18日
无监督元学习表示学习
CreateAMind
27+阅读 · 2019年1月4日
Unsupervised Learning via Meta-Learning
CreateAMind
44+阅读 · 2019年1月3日
disentangled-representation-papers
CreateAMind
26+阅读 · 2018年9月12日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
2+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
4+阅读 · 2009年12月31日
国家自然科学基金
1+阅读 · 2008年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
Arxiv
18+阅读 · 2022年11月21日
Arxiv
15+阅读 · 2021年7月14日
Arxiv
18+阅读 · 2021年6月10日
Arxiv
13+阅读 · 2021年5月3日
VIP会员
最新内容
综述 | 终端智能体:命令行环境中的 AI Agents
专知会员服务
1+阅读 · 8月24日
综述 | 多模态智能体框架基础与前沿
专知会员服务
1+阅读 · 8月24日
《自适应无人机集群网络开发》
专知会员服务
4+阅读 · 8月24日
无人机与作战飞机最佳精确目标定位技术
专知会员服务
3+阅读 · 8月24日
博士论文 | 大动作空间中的在线与离线策略学习
论文 | 全球负责任 AI 指数 2026 方法论
专知会员服务
7+阅读 · 8月23日
超致命战场空间中的战术通信生存能力
专知会员服务
5+阅读 · 8月23日
综述 | 面向机器人的视觉触觉智能
专知会员服务
6+阅读 · 8月22日
论文 | 基础模型驱动具身智能体安全综述
专知会员服务
4+阅读 · 8月22日
相关VIP内容
【AAAI2022】(2.5+1)D时空场景图用于视频问答
专知会员服务
24+阅读 · 2022年2月21日
专知会员服务
16+阅读 · 2021年10月4日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
2+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
4+阅读 · 2009年12月31日
国家自然科学基金
1+阅读 · 2008年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员