Existing video object removal methods excel at inpainting content "behind" the object and correcting appearance-level artifacts such as shadows and reflections. However, when the removed object has more significant interactions, such as collisions with other objects, current models fail to correct them and produce implausible results. We present VOID, a video object removal framework designed to perform physically-plausible inpainting in these complex scenarios. To train the model, we generate a new paired dataset of counterfactual object removals using Kubric and HUMOTO, where removing an object requires altering downstream physical interactions. During inference, a vision-language model identifies regions of the scene affected by the removed object. These regions are then used to guide a video diffusion model that generates physically consistent counterfactual outcomes. Experiments on both synthetic and real data show that our approach better preserves consistent scene dynamics after object removal compared to prior video object removal methods. We hope this framework sheds light on how to make video editing models better simulators of the world through high-level causal reasoning.
翻译:现有视频物体移除方法在修复物体“后方”内容与修正阴影、反射等表层伪影方面表现优异。然而,当被移除物体存在更显著的交互作用(例如与其他物体的碰撞)时,现有模型无法修正这些交互,导致产生不合理的输出结果。我们提出VOID——一个专为复杂场景下物理合理性修复设计的视频物体移除框架。为训练该模型,我们利用Kubric与HUMOTO数据集生成成对的消歧反事实物体移除数据,其中移除物体需同时改变后续的物理交互过程。在推理阶段,视觉语言模型先识别场景中受被移除物体影响的区域,再引导视频扩散模型生成物理一致的反事实输出。实验表明,无论是在合成数据还是真实数据上,我们的方法相较于现有视频物体移除方法更能保持场景动态的一致性。我们期望该框架能揭示如何通过高层因果推理使视频编辑模型成为更好的世界模拟器。