In this paper, we present RStab, a novel framework for video stabilization that integrates 3D multi-frame fusion through volume rendering. Departing from conventional methods, we introduce a 3D multi-frame perspective to generate stabilized images, addressing the challenge of full-frame generation while preserving structure. The core of our approach lies in Stabilized Rendering (SR), a volume rendering module, which extends beyond the image fusion by incorporating feature fusion. The core of our RStab framework lies in Stabilized Rendering (SR), a volume rendering module, fusing multi-frame information in 3D space. Specifically, SR involves warping features and colors from multiple frames by projection, fusing them into descriptors to render the stabilized image. However, the precision of warped information depends on the projection accuracy, a factor significantly influenced by dynamic regions. In response, we introduce the Adaptive Ray Range (ARR) module to integrate depth priors, adaptively defining the sampling range for the projection process. Additionally, we propose Color Correction (CC) assisting geometric constraints with optical flow for accurate color aggregation. Thanks to the three modules, our RStab demonstrates superior performance compared with previous stabilizers in the field of view (FOV), image quality, and video stability across various datasets.
翻译:本文提出RStab,一种通过体绘制集成3D多帧融合的新型视频稳像框架。与常规方法不同,我们引入3D多帧视角生成稳像图像,在保持结构完整性的同时解决全帧生成难题。该框架的核心在于稳定渲染(SR)模块——一种体绘制模块,其突破传统图像融合范畴,通过特征融合实现多帧信息整合。具体而言,SR模块通过投影变换将多帧特征与颜色信息映射至3D空间,融合为描述子以渲染稳像图像。然而,投影精准度直接影响变换信息质量,而动态区域是显著影响因素。为此,我们提出自适应射线范围(ARR)模块,通过集成深度先验自适应定义投影过程的采样范围。此外,我们设计色彩校正(CC)模块,借助光流辅助几何约束实现精准色彩聚合。得益于这三个模块,RStab在视野(FOV)、图像质量及视频稳定性方面均优于现有稳像方法。