Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-language backbones. However, most existing VLAs rely primarily on 2D visual representations, which limit their ability to reason about fine-grained geometry and spatial grounding - capabilities that are essential for precise and robust manipulation in 3D environments. In this paper, we propose PointACT, a dual-system 3D-aware VLA policy that integrates hierarchical 3D point cloud representations directly into the action decoding process. PointACT employs a multi-scale point-action interaction mechanism with efficient bottleneck window self-attention, enabling evolving action tokens to densely attend to both local geometric detail and global scene structure. We evaluate PointACT on the LIBERO and RLBench benchmarks and systematically compare it against monolithic and dual-system VLA baselines, including variants augmented with point cloud inputs. PointACT achieves consistent improvements across both benchmarks, increasing success rates by 10% on the challenging RLBench-10Tasks suite over state-of-the-art pretrained VLAs, with even larger gains when the vision-language backbone is frozen and the action expert is trained from scratch. Extensive ablation studies demonstrate that tightly coupling hierarchical 3D geometry with pretrained 2D semantic representations is critical for robust and spatially grounded robot control. Our results also highlight the promise of pretrained 3D representations for 3D-aware VLA policies.
翻译:摘要:视觉-语言-动作(Vision-Language-Action, VLA)模型通过利用大型预训练视觉-语言骨干网络,在通用机器人操作中展现出强大潜力。然而,现有大多数VLA模型主要依赖二维视觉表征,这限制了其对精细几何结构与空间定位的推理能力——而这些能力对于在三维环境中实现精准且鲁棒的操作至关重要。本文提出PointACT,一种双系统三维感知VLA策略,将分层三维点云表征直接集成至动作解码过程。PointACT采用基于高效瓶颈窗口自注意力的多尺度点-动作交互机制,使得进化中的动作令牌能够同时密集关注局部几何细节与全局场景结构。我们在LIBERO和RLBench基准上评估PointACT,并将其与单系统及双系统VLA基线(包括增强点云输入的变体)进行系统比较。PointACT在两个基准上均实现持续改进,在挑战性更强的RLBench-10Tasks套件上,相较现有最优预训练VLA模型成功率提升10%,当视觉-语言骨干网络被冻结且动作专家从头训练时,性能提升更为显著。大量消融研究表明,将分层三维几何结构与预训练二维语义表征紧密耦合,对于实现鲁棒且具有空间定位能力的机器人控制至关重要。我们的结果还凸显了预训练三维表征在三维感知VLA策略中的应用潜力。