The advent of Chat-GPT has led to a surge of interest in Embodied AI. However, many existing Embodied AI models heavily rely on massive interactions with training environments, which may not be practical in real-world situations. To this end, the Maniskill2 has introduced a full-physics simulation benchmark for manipulating various 3D objects. This benchmark enables agents to be trained using diverse datasets of demonstrations and evaluates their ability to generalize to unseen scenarios in testing environments. In this paper, we propose a novel two-stage fine-tuning strategy that aims to further enhance the generalization capability of our model based on the Maniskill2 benchmark. Through extensive experiments, we demonstrate the effectiveness of our approach by achieving the 1st prize in all three tracks of the ManiSkill2 Challenge. Our findings highlight the potential of our method to improve the generalization abilities of Embodied AI models and pave the way for their ractical applications in real-world scenarios. All codes and models of our solution is available at https://github.com/xtli12/GXU-LIPE.git
翻译:Chat-GPT的出现引发了人们对具身智能的极大兴趣。然而,许多现有的具身智能模型严重依赖与训练环境的大量交互,这在现实场景中可能并不实用。为此,Maniskill2引入了一个全物理仿真基准,用于操作各种三维物体。该基准使智能体能够使用多样化的演示数据集进行训练,并评估其在测试环境中泛化到未见场景的能力。本文提出了一种新颖的两阶段微调策略,旨在基于Maniskill2基准进一步提升模型的泛化能力。通过大量实验,我们在ManiSkill2挑战赛的所有三个赛道中均获得第一名,证明了该方法的有效性。我们的研究结果凸显了该方法在提升具身智能模型泛化能力方面的潜力,并为其在现实场景中的实际应用铺平了道路。我们方案的所有代码和模型均可在https://github.com/xtli12/GXU-LIPE.git获取。