Simulation has become an essential tool for evaluating and improving vision-language-action (VLA) policies, offering scalable, reproducible, and controllable alternatives to costly real-world robot evaluation. Recent simulation benchmarks have made substantial progress on realism and diversity, yet these platforms have not been widely adopted as reliable proxies for real-world policy evaluation. In this work, we investigate this issue through the lens of sim-and-real correlation. We conduct a systematic study across multiple simulation platforms, VLA policies, tasks, and perturbation factors, measuring whether simulated evaluation preserves real-world conclusions in terms of policy ranking consistency, performance correlation, and perturbation-wise failure patterns. This analysis allows us to characterize the limitations of existing simulators and identify what kinds of simulation signals are more aligned with real-world deployment. We further examine how users should exploit simulation for policy improvement, including when simulator-based finetuning is beneficial and how the amount of post-training data affects sim-and-real alignment. Overall, our work provides a unified framework for measuring, interpreting, and improving the usefulness of simulation for VLA policies, offering guidance both for simulator designers and for practitioners who use simulation as part of the policy development pipeline.
翻译:仿真已成为评估和改进视觉-语言-动作(VLA)策略的重要工具,为昂贵的真实世界机器人评估提供了可扩展、可复现且可控的替代方案。近年来,仿真基准平台在真实感和多样性方面取得显著进展,但这些平台尚未被广泛视为真实世界策略评估的可靠代理。本研究从仿真与真实世界相关性的角度探究此问题。我们系统考察了多个仿真平台、VLA策略、任务及扰动因素,衡量仿真评估是否在策略排名一致性、性能相关性及扰动诱导的故障模式上反映真实世界结论。该分析使我们得以刻画现有仿真器的局限性,并识别何种仿真信号更贴近真实世界部署。我们进一步探讨了用户应如何利用仿真改进策略,包括在何种条件下基于仿真的微调有效,以及后训练数据量如何影响仿真与真实世界的对齐。总体而言,本研究提供了一个统一框架,用于衡量、解释和提升仿真对VLA策略的有效性,为仿真器设计者和将仿真纳入策略开发流程的实践者提供指导。