This paper accelerates video perception, such as semantic segmentation and human pose estimation, by levering cross-frame redundancies. Unlike the existing approaches, which avoid redundant computations by warping the past features using optical-flow or by performing sparse convolutions on frame differences, we approach the problem from a new perspective: low-bit quantization. We observe that residuals, as the difference in network activations between two neighboring frames, exhibit properties that make them highly quantizable. Based on this observation, we propose a novel quantization scheme for video networks coined as Residual Quantization. ResQ extends the standard, frame-by-frame, quantization scheme by incorporating temporal dependencies that lead to better performance in terms of accuracy vs. bit-width. Furthermore, we extend our model to dynamically adjust the bit-width proportional to the amount of changes in the video. We demonstrate the superiority of our model, against the standard quantization and existing efficient video perception models, using various architectures on semantic segmentation and human pose estimation benchmarks.
翻译:本文通过利用跨帧冗余性来加速视频感知任务(如语义分割和人体姿态估计)。与现有方法(通过光流扭曲历史特征或对帧差执行稀疏卷积来避免冗余计算)不同,我们从低比特量化这一新视角处理该问题。我们观察到,作为相邻两帧网络激活值之差的残差,具有高度可量化的特性。基于此发现,我们提出一种名为残差量化的视频网络新型量化方案。ResQ通过引入时序依赖关系扩展了标准的逐帧量化方案,从而在精度与比特宽度之间取得更优性能。此外,我们扩展了模型使其能够根据视频中变化量动态调整比特宽度。通过在语义分割和人体姿态估计基准上采用多种架构的对比实验,我们证明了模型相较于标准量化和现有高效视频感知模型的优越性。