Recent progress in video generation has shown impressive visual synthesis capabilities. However, open-domain customized video generation remains limited by the lack of large-scale, annotated datasets capturing diverse identity-specific attributes. To address this, we introduce PexelsCustom-1M, the first publicly available million-scale dataset for identity-preserving video generation, containing one million curated <identity, text, video> triplets across 8,000+ categories. Leveraging this, we propose CustoMDiT, a parameter-efficient framework that adapts a pretrained multimodal Diffusion Transformer into a customized video generator with only 8% additional learnable parameters. Our method surpasses prior state-of-the-art. However, benchmarks such as DreamBooth cover only 100 classes, which is insufficient for real-world applications. To overcome this, we construct OpenCustom, a new benchmark with 1,000+ categories, created via cross-dataset knowledge fusion from ImageNet and MS-COCO. Extensive experiments confirm the advantages of both our dataset and model. We will open-source the entire ecosystem--including dataset, pipeline, benchmark, and implementations--to support further research.
翻译:视频生成领域的最新进展展示了令人瞩目的视觉合成能力。然而,开放域定制化视频生成仍受限于缺乏大规模、带标注的、捕捉多样化身份特定属性的数据集。为解决这一问题,我们推出了PexelsCustom-1M,这是首个公开可用的百万级身份保留视频生成数据集,包含100万条精心整理的<身份,文本,视频>三元组,涵盖8000余个类别。基于此数据集,我们提出了CustoMDiT,这是一种参数高效的框架,通过仅增加8%的可学习参数,将预训练的多模态扩散Transformer适配为定制化视频生成器。我们的方法超越了此前最先进的水平。然而,DreamBooth等基准仅覆盖100个类别,难以满足实际应用需求。为此,我们通过跨数据集知识融合(基于ImageNet和MS-COCO)构建了OpenCustom——一个包含1000多个类别的新基准。大量实验证实了我们的数据集和模型的优势。我们将开源整个生态系统——包括数据集、流水线、基准及实现代码——以支持进一步研究。