It is a long-standing problem in robotics to develop agents capable of executing diverse manipulation tasks from visual observations in unstructured real-world environments. To achieve this goal, the robot needs to have a comprehensive understanding of the 3D structure and semantics of the scene. In this work, we present $\textbf{GNFactor}$, a visual behavior cloning agent for multi-task robotic manipulation with $\textbf{G}$eneralizable $\textbf{N}$eural feature $\textbf{F}$ields. GNFactor jointly optimizes a generalizable neural field (GNF) as a reconstruction module and a Perceiver Transformer as a decision-making module, leveraging a shared deep 3D voxel representation. To incorporate semantics in 3D, the reconstruction module utilizes a vision-language foundation model ($\textit{e.g.}$, Stable Diffusion) to distill rich semantic information into the deep 3D voxel. We evaluate GNFactor on 3 real robot tasks and perform detailed ablations on 10 RLBench tasks with a limited number of demonstrations. We observe a substantial improvement of GNFactor over current state-of-the-art methods in seen and unseen tasks, demonstrating the strong generalization ability of GNFactor. Our project website is https://yanjieze.com/GNFactor/ .
翻译:在非结构化真实环境中,开发能够从视觉观察中执行多样化操作任务的智能体是机器人学领域长期存在的难题。为实现这一目标,机器人需要全面理解场景的三维结构与语义信息。本文提出$\textbf{GNFactor}$——一种基于$\textbf{通}$用$\textbf{神}$经特征$\textbf{场}$的多任务机器人操作视觉行为克隆智能体。GNFactor将通用神经场(GNF)作为重建模块与Perceiver Transformer决策模块联合优化,共享深度三维体素表示。为将语义信息融入三维空间,重建模块利用视觉-语言基础模型(如Stable Diffusion)将丰富语义信息蒸馏至深度三维体素。我们在3项真实机器人任务上评估GNFactor,并在10项RLBench任务上基于少量演示数据进行详细消融实验。实验表明,在可见与未见任务中,GNFactor相比当前最先进方法均有显著提升,展现了其强大的泛化能力。项目网站:https://yanjieze.com/GNFactor/