We propose the NeRF-LEBM, a likelihood-based top-down 3D-aware 2D image generative model that incorporates 3D representation via Neural Radiance Fields (NeRF) and 2D imaging process via differentiable volume rendering. The model represents an image as a rendering process from 3D object to 2D image and is conditioned on some latent variables that account for object characteristics and are assumed to follow informative trainable energy-based prior models. We propose two likelihood-based learning frameworks to train the NeRF-LEBM: (i) maximum likelihood estimation with Markov chain Monte Carlo-based inference and (ii) variational inference with the reparameterization trick. We study our models in the scenarios with both known and unknown camera poses. Experiments on several benchmark datasets demonstrate that the NeRF-LEBM can infer 3D object structures from 2D images, generate 2D images with novel views and objects, learn from incomplete 2D images, and learn from 2D images with known or unknown camera poses.
翻译:我们提出NeRF-LEBM,一种基于似然的由上至下三维感知二维图像生成模型,该模型通过神经辐射场(NeRF)实现三维表示,并通过可微分体渲染实现二维成像过程。该模型将图像表示为从三维物体到二维图像的渲染过程,并以某些潜变量为条件,这些潜变量用于解释物体特征,并假设遵循信息丰富的可训练能量先验模型。我们提出两种基于似然的学习框架来训练NeRF-LEBM:(i)基于马尔可夫链蒙特卡洛推断的最大似然估计,以及(ii)基于重参数化技巧的变分推断。我们在已知和未知相机姿态的场景下研究我们的模型。在多个基准数据集上的实验表明,NeRF-LEBM能够从二维图像推断三维物体结构,生成具有新视角和新物体的二维图像,从不完整的二维图像中学习,以及从已知或未知相机姿态的二维图像中学习。