This paper presents JAWS, an optimization-driven approach that achieves the robust transfer of visual cinematic features from a reference in-the-wild video clip to a newly generated clip. To this end, we rely on an implicit-neural-representation (INR) in a way to compute a clip that shares the same cinematic features as the reference clip. We propose a general formulation of a camera optimization problem in an INR that computes extrinsic and intrinsic camera parameters as well as timing. By leveraging the differentiability of neural representations, we can back-propagate our designed cinematic losses measured on proxy estimators through a NeRF network to the proposed cinematic parameters directly. We also introduce specific enhancements such as guidance maps to improve the overall quality and efficiency. Results display the capacity of our system to replicate well known camera sequences from movies, adapting the framing, camera parameters and timing of the generated video clip to maximize the similarity with the reference clip.
翻译:本文提出JAWS,一种基于优化的方法,能够将参考野生视频片段中的视觉电影化特征稳健地迁移至新生成的视频片段。为此,我们采用隐式神经表示(INR)来计算与参考片段共享相同电影化特征的片段。我们在INR框架中提出了一个通用的摄像机优化问题公式,用于计算内外参以及时序参数。通过利用神经表示的可微分性,我们能够将设计在代理估计器上的电影化损失,通过NeRF网络反向传播至所提出的电影化参数。此外,我们还引入了引导图等特定增强技术,以提升整体效果与效率。实验结果展示了系统复制电影中经典摄像机运动序列的能力,能够调整生成视频片段的取景、摄像机参数及时序,以最大化与参考片段的相似性。