Adaptive video streaming requires efficient bitrate ladder construction to meet heterogeneous network conditions and end-user demands. Per-title optimized encoding typically traverses numerous encoding parameters to search the Pareto-optimal operating points for each video. Recently, researchers have attempted to predict the content-optimized bitrate ladder for pre-encoding overhead reduction. However, existing methods commonly estimate the encoding parameters on the Pareto front and still require subsequent pre-encodings. In this paper, we propose to directly predict the optimal transcoding resolution at each preset bitrate for efficient bitrate ladder construction. We adopt a Temporal Attentive Gated Recurrent Network to capture spatial-temporal features and predict transcoding resolutions as a multi-task classification problem. We demonstrate that content-optimized bitrate ladders can thus be efficiently determined without any pre-encoding. Our method well approximates the ground-truth bitrate-resolution pairs with a slight Bj{\o}ntegaard Delta rate loss of 1.21% and significantly outperforms the state-of-the-art fixed ladder.
翻译:自适应视频流需要高效的码率阶梯构建,以满足异构网络条件和终端用户需求。每标题优化编码通常需要遍历大量编码参数,以搜索每个视频的帕累托最优工作点。近年来,研究人员尝试预测内容优化的码率阶梯,以减少预编码开销。然而,现有方法通常估计帕累托前沿上的编码参数,仍需进行后续预编码。本文提出直接预测每个预设码率下的最优转码分辨率,以实现高效的码率阶梯构建。我们采用时间注意力门控循环网络来捕捉时空特征,并将转码分辨率预测视为多任务分类问题。实验证明,无需任何预编码即可高效确定内容优化的码率阶梯。我们的方法能够较好地逼近真实码率-分辨率对,仅产生1.21%的轻微Bjøntegaard Delta率损失,并显著优于最先进的固定阶梯方法。