Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robotic control, with test-time scaling (TTS) gaining attention to enhance robustness beyond training. However, existing TTS methods for VLAs require additional training, verifiers, and multiple forward passes, making them impractical for deployment. Moreover, they intervene only at action decoding while keeping visual representations fixed-insufficient under perceptual ambiguity, where reconsidering how to perceive is as important as deciding what to do. To address these limitations, we propose SCALE, a simple inference strategy that jointly modulates visual perception and action based on 'self-uncertainty', inspired by uncertainty-driven exploration in Active Inference theory-requiring no additional training, no verifier, and only a single forward pass. SCALE broadens exploration in both perception and action under high uncertainty, while focusing on exploitation when confident-enabling adaptive execution across varying conditions. Experiments on simulated and real-world benchmarks demonstrate that SCALE improves state-of-the-art VLAs and outperforms existing TTS methods while maintaining single-pass efficiency.
翻译:视觉-语言-动作(VLA)模型已成为通用机器人控制领域的范式,其时序扩展推理(TTS)因能增强训练之外的鲁棒性而受到关注。然而,现有VLA的TTS方法需要额外训练、验证器及多次前向传播,导致实际部署困难。更关键的是,这些方法仅在动作解码阶段进行干预,却保持视觉表征固定——这难以应对感知模糊性场景,因为在此类情境中,重新思考感知方式与决定执行动作同等重要。为解决上述局限,我们提出SCALE——一种受主动推理理论中不确定性驱动探索机制启发的简单推理策略,通过基于“自不确定性”联合调控视觉感知与动作生成。该方法无需额外训练、验证器,仅需单次前向传播。在高不确定性条件下,SCALE在感知与动作空间同时扩大探索范围;而在置信度高时则聚焦于利用已有信息——从而实现跨不同场景的自适应执行。仿真与真实世界基准实验表明,SCALE在提升最优VLA模型性能的同时,以单次传播效率优于现有TTS方法。