In this work, we use multi-view aerial images to reconstruct the geometry, lighting, and material of facades using neural signed distance fields (SDFs). Without the requirement of complex equipment, our method only takes simple RGB images captured by a drone as inputs to enable physically based and photorealistic novel-view rendering, relighting, and editing. However, a real-world facade usually has complex appearances ranging from diffuse rocks with subtle details to large-area glass windows with specular reflections, making it hard to attend to everything. As a result, previous methods can preserve the geometry details but fail to reconstruct smooth glass windows or verse vise. In order to address this challenge, we introduce three spatial- and semantic-adaptive optimization strategies, including a semantic regularization approach based on zero-shot segmentation techniques to improve material consistency, a frequency-aware geometry regularization to balance surface smoothness and details in different surfaces, and a visibility probe-based scheme to enable efficient modeling of the local lighting in large-scale outdoor environments. In addition, we capture a real-world facade aerial 3D scanning image set and corresponding point clouds for training and benchmarking. The experiment demonstrates the superior quality of our method on facade holistic inverse rendering, novel view synthesis, and scene editing compared to state-of-the-art baselines.
翻译:本文利用多视角航拍图像,通过神经符号距离场(SDFs)对建筑立面的几何结构、光照与材质进行重建。该方法无需复杂设备,仅需无人机拍摄的RGB图像作为输入,即可实现基于物理的光真实感新视角渲染、重光照及编辑。然而真实世界中的建筑立面通常呈现复杂视觉特征——从具有细微纹理的漫反射岩石到大面积镜面反射玻璃窗,这种多样性使得整体建模极具挑战。现有方法虽能保留几何细节,但难以重建平滑的玻璃窗,反之亦然。为解决该难题,我们提出三项空间与语义自适应优化策略:基于零样本分割技术的语义正则化方法以提升材质一致性、兼顾不同表面平滑度与细节的频率感知几何正则化策略,以及基于可见性探针的大规模户外场景局部光照高效建模方案。此外,我们采集了真实建筑立面的航拍三维扫描图像集及对应点云,用于模型训练与基准测试。实验表明,相较于现有最优基线方法,本方法在建筑立面整体逆渲染、新视角合成与场景编辑任务中均展现出更优越的性能。