While recent text-guided video editing models excel at elementary tasks (e.g., style transfer, object insertion), real-world user requests are highly compositional. A single prompt often demands multiple coupled edits, such as modifying subjects, actions, and camera views, while strictly preserving unrelated spatiotemporal content. Existing benchmarks, heavily constrained by isolated edits and coarse global metrics, fail to diagnose how models handle such complex workflows. To address this gap, we introduce CoVEBench, a compositional video editing benchmark comprising 416 curated source videos, 626 multi-point editing instructions, and 9,990 fine-grained checklist items. Covering diverse editing dimensions, CoVEBench evaluates models via MLLM-judged instruction compliance and video fidelity, alongside automated metrics for video quality. Extensive experiments reveal that compositional editing remains a profound challenge: current models frequently omit edits, violate preservation constraints, or introduce artifacts when handling multiple operations simultaneously. CoVEBench provides a challenging, diagnostic testbed to advance video editing toward realistic user workflows.
翻译:尽管最近的文本引导视频编辑模型在基础任务(如风格迁移、物体插入)上表现出色,但现实用户请求具有高度组合性。单个提示往往要求多重耦合编辑,例如同时修改主体、动作和摄像机视角,同时严格保留无关的时空内容。现有基准测试受限于孤立编辑和粗粒度的全局指标,无法诊断模型如何处理此类复杂工作流。为填补这一空白,我们提出CoVEBench——一个组合式视频编辑基准测试,包含416个精选源视频、626条多点编辑指令及9,990个细粒度检查项。CoVEBench覆盖多种编辑维度,通过基于多模态大语言模型(MLLM)判定的指令遵循度与视频保真度,以及视频质量的自动化指标来评估模型。大量实验表明,组合式编辑仍是一个深刻挑战:当前模型在同时处理多项操作时,常出现编辑遗漏、保真约束违反或引入伪影等问题。CoVEBench为将视频编辑推向更贴近用户真实工作流提供了具有挑战性的诊断性测试平台。