As autonomous systems grow more advanced, objective metrics to evaluate their ethical and legal compliance are critical for informing end users of their limitations and ensuring accountability of those who misuse them. Current ethical embodied AI frameworks remain mostly qualitative, focusing on system design (through safety guardrails or targeted red teaming), and the realized guardrails often directly disallow unsafe behavior without providing the user with an override or interpretable reason. Instead, there is a need for computable metrics through rigorous testing that allow a user to determine the applicability of the system to the task. To address this gap, we introduce the Reference Ethical Benchmark for Autonomy Readiness (REBAR), a quantitative test and evaluation framework for autonomous systems. REBAR maps operating metrics into a computable Autonomy Readiness Level (ARL) rubric that can quantify ethical performance. Key innovations of the framework include a neuro-symbolic Large Language Model (LLM) approach to calculate and explain the ethical difficulty of scenarios, LLM-driven at-scale generation of test instances, and a versatile, photorealistic simulation environment. By evaluating white-box autonomy solutions through this rigorous testing pipeline, REBAR delivers an objective and repeatable benchmark score, bridging the gap between abstract principles and verifiable, accountable autonomy.
翻译:随着自主系统日益先进,量化评估其伦理与法律合规性的客观指标,对于告知最终用户其局限性以及确保滥用者承担责任至关重要。当前的具身AI伦理框架大多仍偏重定性分析,聚焦于系统设计(通过安全护栏或定向红队测试),且已实现的护栏常常直接禁止不安全行为,而不向用户提供覆盖选项或可解释的理由。相反,我们需要通过严格测试获得可计算的指标,从而让用户能够判断系统对特定任务的适用性。为填补这一空白,我们提出了自主系统准备就绪的伦理基准参考(REBAR),这是一个用于自主系统的量化测试与评估框架。REBAR将运行指标映射到一个可计算的自主准备就绪等级(ARL)评分标准,从而量化伦理表现。该框架的关键创新包括采用神经符号大语言模型方法计算并解释场景的伦理难度、利用大语言模型驱动的大规模测试实例生成,以及一个通用的、逼真的仿真环境。通过这一严格测试流水线评估白盒自主解决方案,REBAR提供了客观且可重复的基准评分,架起了抽象原则与可验证、负责任的自主系统之间的桥梁。