Large-scale pretraining has made Vision-Language-Action (VLA) models promising foundations for generalist robot manipulation, yet adapting them to downstream tasks remains necessary. However, the common practice of full fine-tuning treats pretraining as initialization and can shift broad priors toward narrow training-distribution patterns. We propose PriorVLA, a novel framework that preserves pretrained priors and learns to leverage them for effective adaptation. PriorVLA keeps a frozen Prior Expert as a read-only prior source and trains an Adaptation Expert for downstream specialization. Expert Queries capture scene priors from the pretrained VLM and motor priors from the Prior Expert, integrating both into the Adaptation Expert to guide adaptation. Together, PriorVLA updates only 25% of the parameters updated by full fine-tuning. Across RoboTwin 2.0, LIBERO, and real-world tasks, PriorVLA achieves stronger overall performance than full fine-tuning and state-of-the-art VLA baselines, with the largest gains under out-of-distribution (OOD) and few-shot settings. PriorVLA improves over pi0.5 by 11 points on RoboTwin 2.0-Hard and achieves 99.1% average success on LIBERO. Across eight real-world tasks and two embodiments, PriorVLA reaches 81% in-distribution (ID) and 57% OOD success with standard data. With only 10 demonstrations per task, PriorVLA reaches 48% ID and 32% OOD success, surpassing pi0.5 by 24 and 22 points, respectively.
翻译:大规模预训练使视觉-语言-动作(VLA)模型成为通用机器人操控的坚实基础,但将其适配至下游任务仍不可或缺。然而,常规的全参数微调将预训练视为初始化过程,可能将宽泛先验知识窄化为训练分布模式。本文提出PriorVLA——一种新颖的框架,它保持预训练先验知识并学习利用这些知识实现高效适配。PriorVLA冻结先验专家作为只读先验源,同时训练适配专家用于下游任务特化。专家查询从预训练VLM中捕获场景先验、从先验专家中捕获运动先验,并将两者整合至适配专家以指导适配过程。相较于全参数微调,PriorVLA仅更新25%的参数。在RoboTwin 2.0、LIBERO及真实世界任务中,PriorVLA在整体性能上超越全参数微调及最先进的VLA基线,在分布外和少样本场景下提升最为显著。PriorVLA在RoboTwin 2.0-Hard任务上较pi0.5提升11个百分点,在LIBERO上实现99.1%的平均成功率。在涉及两种具身形态的八项真实世界任务中,PriorVLA在标准数据下达到81%的分布内成功率和57%的分布外成功率。当每项任务仅使用10个演示样本时,PriorVLA仍能取得48%的分布内成功率和32%的分布外成功率,分别超越pi0.5达24和22个百分点。