The Agent and AIGC (Artificial Intelligence Generated Content) technologies have recently made significant progress. We propose AesopAgent, an Agent-driven Evolutionary System on Story-to-Video Production. AesopAgent is a practical application of agent technology for multimodal content generation. The system integrates multiple generative capabilities within a unified framework, so that individual users can leverage these modules easily. This innovative system would convert user story proposals into scripts, images, and audio, and then integrate these multimodal contents into videos. Additionally, the animating units (e.g., Gen-2 and Sora) could make the videos more infectious. The AesopAgent system could orchestrate task workflow for video generation, ensuring that the generated video is both rich in content and coherent. This system mainly contains two layers, i.e., the Horizontal Layer and the Utility Layer. In the Horizontal Layer, we introduce a novel RAG-based evolutionary system that optimizes the whole video generation workflow and the steps within the workflow. It continuously evolves and iteratively optimizes workflow by accumulating expert experience and professional knowledge, including optimizing the LLM prompts and utilities usage. The Utility Layer provides multiple utilities, leading to consistent image generation that is visually coherent in terms of composition, characters, and style. Meanwhile, it provides audio and special effects, integrating them into expressive and logically arranged videos. Overall, our AesopAgent achieves state-of-the-art performance compared with many previous works in visual storytelling. Our AesopAgent is designed for convenient service for individual users, which is available on the following page: https://aesopai.github.io/.
翻译:智能体与AIGC(人工智能生成内容)技术近期取得了显著进展。我们提出AesopAgent,一种基于智能体驱动的故事生成视频进化系统。AesopAgent是将智能体技术应用于多模态内容生成的实用范例,该系统将多种生成能力整合于统一框架内,使个体用户能够便捷地调用这些模块。该创新系统可将用户的故事提案转换为脚本、图像和音频,进而将这些多模态内容整合为视频。此外,动画单元(如Gen-2和Sora)可增强视频的感染力。AesopAgent系统能够编排视频生成的任务工作流,确保生成的视频内容既丰富又连贯。该系统主要包含两个层面,即水平层和工具层。在水平层中,我们引入了一种基于RAG的进化系统,以优化整个视频生成工作流及其内部步骤。该系统通过积累专家经验和专业知识(包括优化LLM提示词及工具使用方式)持续进化并迭代优化工作流。工具层提供多种工具,能够生成在构图、角色和风格上视觉一致的连贯图像,同时提供音频和特效,将其整合为富有表现力且逻辑清晰的视频。总体而言,相较于以往的视觉叙事工作,我们的AesopAgent实现了最先进的性能。AesopAgent旨在为个体用户提供便捷服务,可通过以下页面获取:https://aesopai.github.io/。