We introduce a video compression algorithm based on instance-adaptive learning. On each video sequence to be transmitted, we finetune a pretrained compression model. The optimal parameters are transmitted to the receiver along with the latent code. By entropy-coding the parameter updates under a suitable mixture model prior, we ensure that the network parameters can be encoded efficiently. This instance-adaptive compression algorithm is agnostic about the choice of base model and has the potential to improve any neural video codec. On UVG, HEVC, and Xiph datasets, our codec improves the performance of a scale-space flow model by between 21% and 27% BD-rate savings, and that of a state-of-the-art B-frame model by 17 to 20% BD-rate savings. We also demonstrate that instance-adaptive finetuning improves the robustness to domain shift. Finally, our approach reduces the capacity requirements of compression models. We show that it enables a competitive performance even after reducing the network size by 70%.
翻译:我们提出了一种基于实例自适应学习的视频压缩算法。对于每条待传输的视频序列,我们对预训练的压缩模型进行微调。最优参数与潜在码字一同传输至接收端。通过采用合适的混合模型先验对参数更新进行熵编码,我们确保了网络参数能够被高效编码。这种实例自适应压缩算法与基础模型的选择无关,并具有改进任何神经视频编解码器的潜力。在UVG、HEVC和Xiph数据集上,我们的编解码器为尺度空间流模型带来了21%至27%的BD-rate节省,为先进的B帧模型带来了17%至20%的BD-rate节省。我们还证明了实例自适应微调增强了模型对域迁移的鲁棒性。最后,我们的方法降低了压缩模型的容量需求。即使将网络规模减少70%,该方法仍能实现具有竞争力的性能。