Chinese landscape painting is a gem of Chinese cultural and artistic heritage that showcases the splendor of nature through the deep observations and imaginations of its painters. Limited by traditional techniques, these artworks were confined to static imagery in ancient times, leaving the dynamism of landscapes and the subtleties of artistic sentiment to the viewer's imagination. Recently, emerging text-to-video (T2V) diffusion methods have shown significant promise in video generation, providing hope for the creation of dynamic Chinese landscape paintings. However, challenges such as the lack of specific datasets, the intricacy of artistic styles, and the creation of extensive, high-quality videos pose difficulties for these models in generating Chinese landscape painting videos. In this paper, we propose CLV-HD (Chinese Landscape Video-High Definition), a novel T2V dataset for Chinese landscape painting videos, and ConCLVD (Controllable Chinese Landscape Video Diffusion), a T2V model that utilizes Stable Diffusion. Specifically, we present a motion module featuring a dual attention mechanism to capture the dynamic transformations of landscape imageries, alongside a noise adapter to leverage unsupervised contrastive learning in the latent space. Following the generation of keyframes, we employ optical flow for frame interpolation to enhance video smoothness. Our method not only retains the essence of the landscape painting imageries but also achieves dynamic transitions, significantly advancing the field of artistic video generation. The source code and dataset are available at https://anonymous.4open.science/r/ConCLVD-EFE3.
翻译:中国山水画是中华文化艺术瑰宝,通过画家的深度观察与想象展现自然壮美。受传统技法所限,古代此类艺术作品仅能呈现静态意象,山川的动态韵律与艺术情感的微妙变化需由观者自行想象。近年来,新兴的文本到视频(T2V)扩散方法在视频生成领域展现出显著潜力,为动态山水画的创作带来了希望。然而,专用数据集的缺失、艺术风格的复杂性以及高质量长视频的生成难题,给模型创作中国山水画视频带来了挑战。本文提出CLV-HD(中国山水高清视频)——首个面向中国山水画视频的T2V数据集,以及ConCLVD(可控中国山水视频扩散)——基于Stable Diffusion的T2V模型。具体而言,我们设计了具有双重注意力机制的运动模块以捕捉山水意象的动态变换,并结合噪声适配器在潜在空间中利用无监督对比学习。在关键帧生成后,采用光流法进行帧插值以提升视频流畅度。我们的方法既保留了山水画意象的精髓,又实现了动态过渡,显著推动了艺术视频生成领域的发展。源代码与数据集已发布于https://anonymous.4open.science/r/ConCLVD-EFE3。