Neural radiance fields (NeRFs) have exhibited potential in synthesizing high-fidelity views of 3D scenes but the standard training paradigm of NeRF presupposes an equal importance for each image in the training set. This assumption poses a significant challenge for rendering specific views presenting intricate geometries, thereby resulting in suboptimal performance. In this paper, we take a closer look at the implications of the current training paradigm and redesign this for more superior rendering quality by NeRFs. Dividing input views into multiple groups based on their visual similarities and training individual models on each of these groups enables each model to specialize on specific regions without sacrificing speed or efficiency. Subsequently, the knowledge of these specialized models is aggregated into a single entity via a teacher-student distillation paradigm, enabling spatial efficiency for online render-ing. Empirically, we evaluate our novel training framework on two publicly available datasets, namely NeRF synthetic and Tanks&Temples. Our evaluation demonstrates that our DaC training pipeline enhances the rendering quality of a state-of-the-art baseline model while exhibiting convergence to a superior minimum.
翻译:神经辐射场(NeRF)在合成三维场景高保真视图方面展现出潜力,但其标准训练范式预设训练集内每张图像具有同等重要性。这一假设对渲染呈现复杂几何结构的特定视图构成了重大挑战,导致性能次优。本文深入审视当前训练范式的影响,并重新设计该范式以通过NeRF实现更优的渲染质量。根据视觉相似性将输入视图划分为多组,并针对每组分别训练独立模型,使得每个模型能够专精于特定区域而无需牺牲速度或效率。随后,通过教师-学生蒸馏范式将这些专门模型的知识聚合为单一实体,从而实现在线渲染的空间效率。实验方面,我们在两个公开数据集(即NeRF合成数据集和Tanks&Temples数据集)上评估了所提新颖训练框架。评估结果表明,我们的DaC训练流程在提升最先进基线模型渲染质量的同时,展现出收敛至更优极小值的能力。