Reference-based video object segmentation is an emerging topic which aims to segment the corresponding target object in each video frame referred by a given reference, such as a language expression or a photo mask. However, language expressions can sometimes be vague in conveying an intended concept and ambiguous when similar objects in one frame are hard to distinguish by language. Meanwhile, photo masks are costly to annotate and less practical to provide in a real application. This paper introduces a new task of sketch-based video object segmentation, an associated benchmark, and a strong baseline. Our benchmark includes three datasets, Sketch-DAVIS16, Sketch-DAVIS17 and Sketch-YouTube-VOS, which exploit human-drawn sketches as an informative yet low-cost reference for video object segmentation. We take advantage of STCN, a popular baseline of semi-supervised VOS task, and evaluate what the most effective design for incorporating a sketch reference is. Experimental results show sketch is more effective yet annotation-efficient than other references, such as photo masks, language and scribble.
翻译:参考式视频对象分割是一个新兴的研究课题,旨在依据给定的参考(如语言描述或照片掩码)对视频每一帧中的对应目标对象进行分割。然而,语言描述有时在传达特定概念时不够明确,且当同一帧中相似对象难以通过语言区分时会产生歧义。同时,照片掩码的标注成本较高,在实际应用中较难提供。本文提出了一项新任务——基于草图的视频对象分割,并构建了相关基准与强大的基线模型。该基准包含三个数据集:Sketch-DAVIS16、Sketch-DAVIS17和Sketch-YouTube-VOS,这些数据集利用手绘草图作为信息丰富且低成本的视频对象分割参考。我们借鉴了半监督视频对象分割任务中常用的基线模型STCN,并评估了整合草图参考的最有效设计方案。实验结果表明,相较于照片掩码、语言描述和涂鸦等其他参考方式,草图参考在分割效果上更优且标注效率更高。