Volumetric video has emerged as a prominent medium within the realm of eXtended Reality (XR) with the advancements in computer graphics and depth capture hardware. Users can fully immersive themselves in volumetric video with the ability to switch their viewport in six degree-of-freedom (DOF), including three rotational dimensions (yaw, pitch, roll) and three translational dimensions (X, Y, Z). Different from traditional 2D videos that are composed of pixel matrices, volumetric videos employ point clouds, meshes, or voxels to represent a volumetric scene, resulting in significantly larger data sizes. While previous works have successfully achieved volumetric video streaming in video-on-demand scenarios, the live streaming of volumetric video remains an unresolved challenge due to the limited network bandwidth and stringent latency constraints. In this paper, we for the first time propose a holistic live volumetric video streaming system, LiveVV, which achieves multi-view capture, scene segmentation \& reuse, adaptive transmission, and rendering. LiveVV contains multiple lightweight volumetric video capture modules that are capable of being deployed without prior preparation. To reduce bandwidth consumption, LiveVV processes static and dynamic volumetric content separately by reusing static data with low disparity and decimating data with low visual saliency. Besides, to deal with network fluctuation, LiveVV integrates a volumetric video adaptive bitrate streaming algorithm (VABR) to enable fluent playback with the maximum quality of experience. Extensive real-world experiment shows that LiveVV can achieve live volumetric video streaming at a frame rate of 24 fps with a latency of less than 350ms.
翻译:体积视频随着计算机图形学与深度采集硬件的进步,已成为扩展现实领域中的重要媒介。用户能够以六自由度切换视口,包括三个旋转维度(偏航角、俯仰角、翻滚角)和三个平移维度(X、Y、Z坐标),从而完全沉浸于体积视频中。不同于由像素矩阵构成的传统二维视频,体积视频采用点云、网格或体素来表达立体场景,导致数据量显著增大。尽管既有研究已成功实现视频点播场景下的体积视频流传输,但受限于网络带宽不足和严格的延迟约束,体积视频的实时直播仍是一个未解决的挑战。本文首次提出一套完整的实时体积视频流媒体系统LiveVV,实现了多视角采集、场景分割与复用、自适应传输及渲染。LiveVV包含多个轻量级体积视频采集模块,可在无需预先准备的情况下部署。为降低带宽消耗,系统通过复用低视差静态数据并剔除低视觉显著性数据的方式,对静态与动态体积内容进行差异化处理。此外,为应对网络波动,LiveVV集成了体积视频自适应码率流算法,以最大体验质量实现流畅播放。大规模真实环境实验表明,LiveVV能以24帧/秒的帧率实现小于350毫秒延迟的实时体积视频流传输。