While diffusion models generate high-fidelity video clips, transforming them into coherent storytelling engines remains challenging. Current agentic pipelines automate this via chained modules but suffer from semantic drift and cascading failures due to independent, handcrafted prompting. We present Co-Director, a hierarchical multi-agent framework formalizing video storytelling as a global optimization problem. To ensure semantic coherence, we introduce hierarchical parameterization: a multi-armed bandit globally identifies promising creative directions, while a local multimodal self-refinement loop mitigates identity drift and ensures sequence-level consistency. This balances the exploration of novel narrative strategies with the exploitation of effective creative configurations. For evaluation, we introduce GenAD-Bench, a 400-scenario dataset of fictional products for personalized advertising. Experiments demonstrate that Co-Director significantly outperforms state-of-the-art baselines, offering a principled approach that seamlessly generalizes to broader cinematic narratives. Project Page: https://co-director-agent.github.io/
翻译:尽管扩散模型能够生成高保真视频片段,但将其转化为连贯的叙事引擎仍具挑战性。当前的智能体流水线通过链式模块实现自动化,但由于依赖独立的手工提示策略,易导致语义漂移和级联失败。我们提出Co-Director,一个将视频叙事形式化为全局优化问题的分层多智能体框架。为确保语义一致性,我们引入分层参数化:多臂老虎机机制全局识别有前途的创意方向,而局部多模态自优化循环缓解身份漂移并保证序列级一致性。该方法平衡了新型叙事策略的探索与有效创意配置的利用。为进行评估,我们构建了包含400个虚构产品广告场景的GenAD-Bench数据集。实验表明,Co-Director显著优于现有基线方法,其原理性设计可无缝泛化至更广泛的电影叙事。项目页面:https://co-director-agent.github.io/