We propose CARFF: Conditional Auto-encoded Radiance Field for 3D Scene Forecasting, a method for predicting future 3D scenes given past observations, such as 2D ego-centric images. Our method maps an image to a distribution over plausible 3D latent scene configurations using a probabilistic encoder, and predicts the evolution of the hypothesized scenes through time. Our latent scene representation conditions a global Neural Radiance Field (NeRF) to represent a 3D scene model, which enables explainable predictions and straightforward downstream applications. This approach extends beyond previous neural rendering work by considering complex scenarios of uncertainty in environmental states and dynamics. We employ a two-stage training of Pose-Conditional-VAE and NeRF to learn 3D representations. Additionally, we auto-regressively predict latent scene representations as a partially observable Markov decision process, utilizing a mixture density network. We demonstrate the utility of our method in realistic scenarios using the CARLA driving simulator, where CARFF can be used to enable efficient trajectory and contingency planning in complex multi-agent autonomous driving scenarios involving visual occlusions.
翻译:我们提出CARFF:用于3D场景预测的条件自编码辐射场(CARFF),一种根据过去观测(如2D第一人称图像)预测未来3D场景的方法。该方法通过概率编码器将图像映射到可能的3D潜在场景配置分布上,并预测假设场景随时间演变的过程。我们的潜在场景表示通过条件全局神经辐射场(NeRF)构建3D场景模型,从而支持可解释的预测及直接的后续应用。该工作超越了先前的神经渲染研究,考虑了环境状态与动态变化中的复杂不确定性场景。我们采用姿态条件变分自编码器(Pose-Conditional-VAE)和NeRF的两阶段训练来学习3D表示。此外,我们利用混合密度网络(MDN)将潜在场景表示的自动回归预测建模为部分可观测马尔可夫决策过程(POMDP)。我们通过CARLA驾驶模拟器在现实场景中验证了该方法——在涉及视觉遮挡的复杂多智能体自动驾驶场景中,CARFF可实现高效的轨迹规划与应急规划。