The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck due to rapid performance saturation and a lack of true generalization testing. Existing benchmarks often exhibit significant domain overlap between training and evaluation, trivializing success rates and obscuring insights into robustness. We introduce RoboLab, a simulation benchmarking framework designed to address these challenges. Concretely, our framework is designed to answer two questions: (1) to what extent can we understand the performance of a real-world policy by analyzing its behavior in simulation, and (2) which factor most strongly affect policy behavior. First, RoboLab enables human-authored and LLM-enabled generation of scenes and tasks in a robot- and policy-agnostic manner within a high-fidelity simulation environment. We introduce an accompanying RoboLab-120 benchmark, consisting of 120 tasks categorized into three competency axes: visual, procedural, relational, across three difficulty levels. Second, we introduce a systematic analysis of real-world policies that quantify both their performance and the sensitivity of their behavior to controlled perturbations, exposing significant performance gap in current state-of-the-art models. By providing granular metrics and a scalable toolset, RoboLab offers a scalable framework for evaluating the true generalization capabilities of task-generalist robotic policies. Project website: https://research.nvidia.com/labs/srl/projects/robolab/.


翻译:通用机器人技术的发展催生了令人瞩目的基础模型,然而由于性能饱和迅速且缺乏真正的泛化测试,基于仿真的基准测试仍是关键瓶颈。现有基准在训练与评估任务之间常呈现显著领域重叠,导致成功率失去区分度且难以揭示对鲁棒性的深入理解。为此,我们提出RoboLab——一种专为攻克上述挑战而设计的仿真基准框架。具体而言,本框架旨在回答两个核心问题:(1)通过分析策略在仿真中的行为,能在多大程度上理解其在真实世界中的性能?(2)何种因素对策略行为的影响最为显著?首先,RoboLab在高保真仿真环境中支持以机器人无关与策略无关的方式,实现人工编写与大语言模型辅助的场景与任务生成。我们引入配套基准RoboLab-120,其包含120项任务,按视觉、程序、关系三大能力维度分为三个难度等级。其次,我们对真实世界策略开展系统性分析,量化其性能表现及行为对受控扰动的敏感度,揭示当前最先进模型存在的显著性能差距。通过提供细粒度指标与可扩展工具集,RoboLab为评估任务通用机器人策略的真实泛化能力提供了可扩展框架。项目网址:https://research.nvidia.com/labs/srl/projects/robolab/。

0
下载
关闭预览

相关内容

军事决策大语言模型综合评价基准
专知会员服务
20+阅读 · 4月1日
【CMU博士论文】利用信息论工具进行基础模型分析
专知会员服务
19+阅读 · 2025年8月31日
【MIT博士论文】理解与提升机器学习模型的表征鲁棒性
专知会员服务
30+阅读 · 2024年8月26日
【斯坦福博士论文】基础模型的数据分布视角,321页pdf
专知会员服务
42+阅读 · 2024年7月8日
NLG任务评价指标BLEU与ROUGE
AINLP
21+阅读 · 2020年5月25日
基于数据的分布式鲁棒优化算法及其应用【附PPT与视频资料】
人工智能前沿讲习班
27+阅读 · 2018年12月13日
Maplab:研究视觉惯性建图和定位的开源框架
泡泡机器人SLAM
16+阅读 · 2018年4月4日
FCS 论坛 | 孟德宇:误差建模原理
FCS
15+阅读 · 2017年8月17日
基于机器学习的KPI自动化异常检测系统
运维帮
13+阅读 · 2017年8月16日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
16+阅读 · 2013年12月31日
VIP会员
最新内容
分层反无人机系统发展新趋势
专知会员服务
7+阅读 · 9月3日
何为协作武器?
专知会员服务
10+阅读 · 9月1日
《理解认知战:超越信息》
专知会员服务
13+阅读 · 9月1日
美国战争部在GenAI.mil上推出OpenAI的ChatGPT Mil
专知会员服务
8+阅读 · 8月31日
人工智能赋能军事维护:重新定义国防战备
专知会员服务
5+阅读 · 8月31日
《美陆军野战手册(2026年):特种部队》
专知会员服务
9+阅读 · 8月31日
相关VIP内容
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
16+阅读 · 2013年12月31日
Top
微信扫码咨询专知VIP会员