In robot imitation learning, influence functions provide a principled approach to quantify each demonstration's effect on robot task outcomes, yet scaling them to billion-parameter Vision-Language-Action (VLA) models is limited by computational and multitask bottlenecks. To this end, we propose ATHENA, an influence function framework tailored for multitask VLA data curation at a billion-parameter scale. Concretely, it leverages the Kronecker structure of linear-layer gradients to reduce projection cost, and approximates dense Hessian inversion with a rank-r Random Truncated Approximation, achieving about a 313.4x speedup in influence computation. Furthermore, ATHENA formulates global and local interactive influence to balance data curation across 50 jointly trained tasks. Extensive evaluations on RoboTwin 2.0 and real-robot deployment, covering 9.34 and 6.90 hours of demonstrations, respectively, show that ATHENA matches or exceeds full-data joint fine-tuning using only 50% of demonstrations in simulation and 66.7% of data across six real-robot tasks. Overall, ATHENA demonstrates its effectiveness for data curation in billion-parameter multitask VLA fine-tuning.
翻译:在机器人模仿学习中,影响函数提供了一种原则性方法,用于量化每个示范对机器人任务结果的影响,但将其扩展到数十亿参数的视觉-语言-动作(VLA)模型受到计算和多任务瓶颈的限制。为此,我们提出ATHENA,一种专为数十亿参数规模的多任务VLA数据筛选定制的影响函数框架。具体而言,它利用线性层梯度的克罗内克结构降低投影成本,并通过秩为r的随机截断近似来逼近稠密海森矩阵求逆,在影响计算中实现了约313.4倍的加速。此外,ATHENA构建了全局与局部交互影响,以平衡跨50个联合训练任务的数据筛选。在RoboTwin 2.0和真实机器人部署上的广泛评估(分别涵盖9.34小时和6.90小时的示范数据)表明,ATHENA在仿真中仅使用50%的示范数据、在六项真实机器人任务中仅使用66.7%的数据即能达到或超越全数据联合微调的性能。总体而言,ATHENA在十亿参数级多任务VLA微调中展示了其在数据筛选方面的有效性。