This paper considers the problem of evaluating an autonomous system's competency in performing a task, particularly when working in dynamic and uncertain environments. The inherent opacity of machine learning models, from the perspective of the user, often described as a `black box', poses a challenge. To overcome this, we propose using a measure called the Surprise index, which leverages available measurement data to quantify whether the dynamic system performs as expected. We show that the surprise index can be computed in closed form for dynamic systems when observed evidence in a probabilistic model if the joint distribution for that evidence follows a multivariate Gaussian marginal distribution. We then apply it to a nonlinear spacecraft maneuver problem, where actions are chosen by a reinforcement learning agent and show it can indicate how well the trajectory follows the required orbit.
翻译:本文研究在动态且不确定环境下评估自主系统执行任务能力的问题。机器学习模型从用户视角来看往往存在内在不透明性,常被称为"黑箱",这给评估带来了挑战。为克服这一困难,我们提出采用"惊奇指数"这一度量方法,通过可利用的测量数据来量化动态系统的实际表现是否符合预期。研究表明:当观测证据在概率模型中的联合分布服从多元高斯边际分布时,惊奇指数可对动态系统进行闭式计算。随后我们将该方法应用于非线性航天器机动问题(其中动作选择由强化学习智能体实现),实验表明该指数能有效反映轨迹对目标轨道的拟合程度。