In the paradigm of AI-generated content (AIGC), there has been increasing attention to transferring knowledge from pre-trained text-to-image (T2I) models to text-to-video (T2V) generation. Despite their effectiveness, these frameworks face challenges in maintaining consistent narratives and handling shifts in scene composition or object placement from a single abstract user prompt. Exploring the ability of large language models (LLMs) to generate time-dependent, frame-by-frame prompts, this paper introduces a new framework, dubbed DirecT2V. DirecT2V leverages instruction-tuned LLMs as directors, enabling the inclusion of time-varying content and facilitating consistent video generation. To maintain temporal consistency and prevent mapping the value to a different object, we equip a diffusion model with a novel value mapping method and dual-softmax filtering, which do not require any additional training. The experimental results validate the effectiveness of our framework in producing visually coherent and storyful videos from abstract user prompts, successfully addressing the challenges of zero-shot video generation.
翻译:在人工智能生成内容(AIGC)的范式下,将预训练文本到图像(T2I)模型的知识迁移至文本到视频(T2V)生成已引起越来越多的关注。尽管这些框架效果显著,但面对单一抽象用户提示词时,它们在维持叙事连贯性以及处理场景构成或物体位置变化方面仍面临挑战。本文探索利用大语言模型(LLM)生成随时间变化的逐帧提示词的能力,提出了一种名为DirecT2V的新框架。DirecT2V将经过指令微调的大语言模型作为导演,使其能够包含随时间变化的内容,并促进连贯的视频生成。为保持时间一致性、防止将映射值赋予不同物体,我们为扩散模型配备了一种新颖的值映射方法以及双softmax滤波技术,且该方法无需额外训练。实验结果验证了我们的框架能够根据抽象用户提示词生成视觉连贯且富有故事性的视频,成功解决了零样本视频生成中的挑战。