While video segmentation has advanced rapidly on short clips and closed-set benchmarks, open-world video segmentation remains largely unexplored. The challenge is twofold: (1) existing methods are not designed to support object discovery and identity maintenance in long videos of dynamic ego-motion, and (2) existing evaluation protocols rely on a rigid 1:1 matching that unfairly penalizes semantically valid predictions with mismatched granularity. To address both gaps, we introduce Savvy, a practical and strong system for zero-shot open-world long-horizon video segmentation. Savvy combines hierarchical mask discovery, deferred admission, and track consolidation to support persistent object discovery, safe track promotion, and stable long-range identity maintenance. We further propose OGA, a granularity-aware evaluation suite for open-world video segmentation. Built on a Granularity-Agnostic (GA) matching protocol, OGA relaxes conventional 1:1 matching to an n:1 mapping, but still enforces temporal rigor by detecting support discontinuities through sever points and scoring each reference object through its dominant coherent fragment. This prevents fragmented or flickering support from being over-rewarded while enabling GA-adapted metrics and structural diagnostics: identity persistence (IP), and identity concentration (IC). On VIPSeg, we show that standard 1:1 evaluation substantially underestimates open-world methods, whereas GA evaluation recovers much of their suppressed performance. On the more realistic long-horizon benchmarks: ScanNet and HM3D, Savvy consistently outperforms strong baselines across both classical and proposed metrics, including STQ, VPQ$_\infty$, IP and IC. Together, these results establish a practical benchmark and a strong baseline for open-world long-horizon video segmentation.
翻译:尽管视频分割在短片段和封闭集基准上取得了快速进展,但开放世界视频分割仍基本未被探索。其挑战来源于两方面:(1) 现有方法并非旨在支持动态自我运动长视频中的物体发现与身份维护;(2) 现有评估协议依赖严格的1:1匹配,不公正地惩罚了具有不匹配粒度的语义有效预测。为填补这两项空白,我们提出Savvy——一个实用且强大的零样本开放世界长时域视频分割系统。Savvy融合了层次化掩码发现、延迟确认和轨迹整合,以支持持续物体发现、安全轨迹提升和稳定长距离身份维护。我们进一步提出OGA——一个面向开放世界视频分割的粒度感知评估套件。OGA基于粒度无关(GA)匹配协议构建,将传统的1:1匹配放宽为n:1映射,同时通过切分点检测支持不连续性、基于主导连贯片段对每个参考物体进行评分,从而保持时域严谨性。这防止了碎片化或闪烁性支持被过度奖励,同时支持GA自适应指标与结构性诊断:身份持久性(IP)和身份集中度(IC)。在VIPSeg上,我们表明标准1:1评估显著低估了开放世界方法,而GA评估则恢复了其被抑制的大部分性能。在更真实的长时域基准ScanNet和HM3D上,Savvy在经典指标与新提出指标(包括STQ、VPQ$_\infty$、IP和IC)上均持续优于强基线。这些成果共同为开放世界长时域视频分割建立了实用基准与强基线。