While video segmentation has advanced rapidly on short clips and closed-set benchmarks, open-world video segmentation remains largely unexplored. The challenge is twofold: (1) existing methods are not designed to support object discovery and identity maintenance in long videos of dynamic ego-motion, and (2) existing evaluation protocols rely on a rigid 1:1 matching that unfairly penalizes semantically valid predictions with mismatched granularity. To address both gaps, we introduce Savvy, a practical and strong system for zero-shot open-world long-horizon video segmentation. Savvy combines hierarchical mask discovery, deferred admission, and track consolidation to support persistent object discovery, safe track promotion, and stable long-range identity maintenance. We further propose OGA, a granularity-aware evaluation suite for open-world video segmentation. Built on a Granularity-Agnostic (GA) matching protocol, OGA relaxes conventional 1:1 matching to an n:1 mapping, but still enforces temporal rigor by detecting support discontinuities through sever points and scoring each reference object through its dominant coherent fragment. This prevents fragmented or flickering support from being over-rewarded while enabling GA-adapted metrics and structural diagnostics: identity persistence (IP), and identity concentration (IC). On VIPSeg, we show that standard 1:1 evaluation substantially underestimates open-world methods, whereas GA evaluation recovers much of their suppressed performance. On the more realistic long-horizon benchmarks: ScanNet and HM3D, Savvy consistently outperforms strong baselines across both classical and proposed metrics, including STQ, VPQ$_\infty$, IP and IC. Together, these results establish a practical benchmark and a strong baseline for open-world long-horizon video segmentation.


翻译:尽管视频分割在短片段和封闭集基准上取得了快速进展,但开放世界视频分割仍基本未被探索。其挑战来源于两方面:(1) 现有方法并非旨在支持动态自我运动长视频中的物体发现与身份维护;(2) 现有评估协议依赖严格的1:1匹配,不公正地惩罚了具有不匹配粒度的语义有效预测。为填补这两项空白,我们提出Savvy——一个实用且强大的零样本开放世界长时域视频分割系统。Savvy融合了层次化掩码发现、延迟确认和轨迹整合,以支持持续物体发现、安全轨迹提升和稳定长距离身份维护。我们进一步提出OGA——一个面向开放世界视频分割的粒度感知评估套件。OGA基于粒度无关(GA)匹配协议构建,将传统的1:1匹配放宽为n:1映射,同时通过切分点检测支持不连续性、基于主导连贯片段对每个参考物体进行评分,从而保持时域严谨性。这防止了碎片化或闪烁性支持被过度奖励,同时支持GA自适应指标与结构性诊断:身份持久性(IP)和身份集中度(IC)。在VIPSeg上,我们表明标准1:1评估显著低估了开放世界方法,而GA评估则恢复了其被抑制的大部分性能。在更真实的长时域基准ScanNet和HM3D上,Savvy在经典指标与新提出指标(包括STQ、VPQ$_\infty$、IP和IC)上均持续优于强基线。这些成果共同为开放世界长时域视频分割建立了实用基准与强基线。

0
下载
关闭预览

相关内容

【ECCV2024】开放世界动态提示与持续视觉表征学习
专知会员服务
25+阅读 · 2024年9月10日
专知会员服务
23+阅读 · 2021年7月5日
【CVPR2021】基于Transformer的视频分割领域
专知会员服务
38+阅读 · 2021年4月16日
【Google】多模态Transformer视频检索,Multi-modal Transformer
专知会员服务
103+阅读 · 2020年7月22日
全景分割这一年,端到端之路
机器之心
14+阅读 · 2018年12月24日
超像素、语义分割、实例分割、全景分割 傻傻分不清?
计算机视觉life
19+阅读 · 2018年11月27日
语义分割+视频分割开源代码集合
极市平台
35+阅读 · 2018年3月5日
一文带你入门视频目标分割(附数据集)
THU数据派
19+阅读 · 2017年10月10日
入门 | 一文概览视频目标分割
机器之心
10+阅读 · 2017年10月6日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 6月17日
VIP会员
最新内容
非对称防御中的自组织临界性:俄乌战争
专知会员服务
1+阅读 · 今天14:36
《战争中的大语言模型监管》
专知会员服务
2+阅读 · 今天14:26
边缘计算的军事应用
专知会员服务
8+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
9+阅读 · 8月8日
相关VIP内容
【ECCV2024】开放世界动态提示与持续视觉表征学习
专知会员服务
25+阅读 · 2024年9月10日
专知会员服务
23+阅读 · 2021年7月5日
【CVPR2021】基于Transformer的视频分割领域
专知会员服务
38+阅读 · 2021年4月16日
【Google】多模态Transformer视频检索,Multi-modal Transformer
专知会员服务
103+阅读 · 2020年7月22日
相关基金
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员