Existing implicit neural representation (INR) methods do not fully exploit spatiotemporal redundancies in videos. Index-based INRs ignore the content-specific spatial features and hybrid INRs ignore the contextual dependency on adjacent frames, leading to poor modeling capability for scenes with large motion or dynamics. We analyze this limitation from the perspective of function fitting and reveal the importance of frame difference. To use explicit motion information, we propose Difference Neural Representation for Videos (DNeRV), which consists of two streams for content and frame difference. We also introduce a collaborative content unit for effective feature fusion. We test DNeRV for video compression, inpainting, and interpolation. DNeRV achieves competitive results against the state-of-the-art neural compression approaches and outperforms existing implicit methods on downstream inpainting and interpolation for $960 \times 1920$ videos.
翻译:现有隐式神经表示(INR)方法未能充分利用视频中的时空冗余。基于索引的INR忽略了内容特定的空间特征,而混合INR忽略了相邻帧的上下文依赖性,导致对具有大运动或动态场景的建模能力不足。我们从函数拟合的角度分析了这一局限性,揭示了帧间差分的重要性。为利用显式运动信息,我们提出了视频差分神经表示(DNeRV),该方法包含内容流与帧差分流两个分支。同时引入协作内容单元以实现高效特征融合。我们在视频压缩、修复和插值任务上测试了DNeRV。实验结果表明,DNeRV在与最先进神经压缩方法的竞争中取得了有竞争力的结果,并在$960 \times 1920$视频的下游修复与插值任务中优于现有隐式方法。