The advent of ChatGPT has introduced innovative methods for information gathering and analysis. However, the information provided by ChatGPT is limited to text, and the visualization of this information remains constrained. Previous research has explored zero-shot text-to-video (TTV) approaches to transform text into videos. However, these methods lacked control over the identity of the generated audio, i.e., not identity-agnostic, hindering their effectiveness. To address this limitation, we propose a novel two-stage framework for person-agnostic video cloning, specifically focusing on TTV generation. In the first stage, we leverage pretrained zero-shot models to achieve text-to-speech (TTS) conversion. In the second stage, an audio-driven talking head generation method is employed to produce compelling videos privided the audio generated in the first stage. This paper presents a comparative analysis of different TTS and audio-driven talking head generation methods, identifying the most promising approach for future research and development. Some audio and videos samples can be found in the following link: https://github.com/ZhichaoWang970201/Text-to-Video/tree/main.
翻译:ChatGPT的出现为信息收集与分析引入了创新方法。然而,ChatGPT提供的信息仅限于文本,这些信息的可视化仍然受到限制。已有研究探索了零样本文本到视频(TTV)方法,试图将文本转换为视频。但这些方法缺乏对生成音频身份的控制,即无法实现身份无关性,从而限制了其有效性。为解决这一局限,我们提出了一种新颖的人无关视频克隆二阶段框架,特别聚焦于TTV生成。在第一阶段,我们利用预训练的零样本模型实现文本到语音(TTS)转换。在第二阶段,采用基于音频驱动的说话头生成方法,利用第一阶段生成的音频生成引人入胜的视频。本文对不同TTS与音频驱动说话头生成方法进行了比较分析,确定了未来研究与开发中最具前景的技术路径。部分音频与视频示例可见于以下链接:https://github.com/ZhichaoWang970201/Text-to-Video/tree/main。