Learning generative object models from unlabelled videos is a long standing problem and required for causal scene modeling. We decompose this problem into three easier subtasks, and provide candidate solutions for each of them. Inspired by the Common Fate Principle of Gestalt Psychology, we first extract (noisy) masks of moving objects via unsupervised motion segmentation. Second, generative models are trained on the masks of the background and the moving objects, respectively. Third, background and foreground models are combined in a conditional "dead leaves" scene model to sample novel scene configurations where occlusions and depth layering arise naturally. To evaluate the individual stages, we introduce the Fishbowl dataset positioned between complex real-world scenes and common object-centric benchmarks of simplistic objects. We show that our approach allows learning generative models that generalize beyond the occlusions present in the input videos, and represent scenes in a modular fashion that allows sampling plausible scenes outside the training distribution by permitting, for instance, object numbers or densities not observed in the training set.
翻译:从无标注视频中学习生成式对象模型是一个长期存在的问题,也是因果场景建模的必要条件。我们将该问题分解为三个较简单的子任务,并为每个子任务提供候选解决方案。受格式塔心理学"共同命运原则"启发,我们首先通过无监督运动分割提取运动对象的(含噪)掩码。其次,分别对背景和运动对象的掩码训练生成式模型。第三,将背景与前景模型以条件式"枯叶"场景模型相结合,以采样新颖的场景配置,使遮挡和深度分层自然产生。为评估各阶段表现,我们引入了介于复杂真实场景与简单物体中心基准之间的Fishbowl数据集。实验表明,我们的方法能够学习到可泛化至输入视频中未出现遮挡情况的生成式模型,并以模块化方式表示场景,通过允许训练集未观测到的对象数量或密度等参数,生成训练分布之外的合理场景。