Multi-shot long-form video generation remains challenging due to identity drift and compounding inconsistencies across shots. While storyboard-driven pipelines improve controllability, they are often executed in a feed-forward manner, with limited mechanisms to incorporate generated visual evidence back into subsequent conditioning. We propose CoTriSyGen, an agentic framework that formulates multi-shot long video generation as a closed-loop visual-text-memory synergy process, where planned intent, persistent memory, and generated visuals are jointly leveraged for iterative correction and long-range coherence. A vision-language-model-based analyzer reasons over this triplet and produces updates to both prompts and memory along two pathways: (i) intra-shot refinement, which triggers targeted regeneration when semantic or compositional violations are detected and refines image-to-video prompt for coherent motions; and (ii) inter-shot refinement, which rewrites subsequent-shot prompts to propagate newly manifested entities or attributes and improve prompt quality (e.g., compositional grounding and cinematic fluency) based on generated evidence. The loop is grounded in an entity-centric memory modeled as a mutable visual state that evolves as the story progresses, which is continuously updated by both the generator and the analyzer by adding new and evolved entities to reflect appearance changes, accumulated multi-view evidence, and multi-entity compositions. Experiments on our curated StoryBench benchmark demonstrate substantial improvements in cross-shot consistency, prompt adherence, and cinematic continuity over representative methods.
翻译:多镜头长视频生成因镜头间身份偏移和复合不一致性仍具挑战性。基于故事板驱动的流程虽能提升可控性,但通常以前馈方式执行,缺乏将生成视觉证据融入后续条件机制的机制。我们提出CoTriSyGen,一种将多镜头长视频生成构建为闭环视觉-文本-记忆协同过程的智能体框架,其中计划意图、持久记忆和生成视觉被联合用于迭代校正和长程连贯性。基于视觉语言模型的分析器对此三联体进行推理,并通过两条路径更新提示词和记忆:(i) 镜头内优化——当检测到语义或构成违规时触发定向再生,并优化图像到视频的提示词以实现连贯运动;(ii) 镜头间优化——基于生成证据重写后续镜头提示词以传播新出现的实体或属性,提升提示词质量(如构成基础和电影流畅度)。该循环以实体中心记忆为基础,建模为随故事进展而演变的可变视觉状态,通过生成器和分析器持续添加新实体及进化实体以反映外观变化、累积多视角证据和多实体构成。在我们自建的StoryBench基准测试上的实验表明,该方法在跨镜头一致性、提示词遵循度和电影连续性上均显著优于代表性方法。