Rapid advancements in artificial intelligence (AI) technology have brought about a plethora of new challenges in terms of governance and regulation. AI systems are being integrated into various industries and sectors, creating a demand from decision-makers to possess a comprehensive and nuanced understanding of the capabilities and limitations of these systems. One critical aspect of this demand is the ability to explain the results of machine learning models, which is crucial to promoting transparency and trust in AI systems, as well as fundamental in helping machine learning models to be trained ethically. In this paper, we present novel quantitative metrics frameworks for interpreting the predictions of classifier and regressor models. The proposed metrics are model agnostic and are defined in order to be able to quantify: i. the interpretability factors based on global and local feature importance distributions; ii. the variability of feature impact on the model output; and iii. the complexity of feature interactions within model decisions. We employ publicly available datasets to apply our proposed metrics to various machine learning models focused on predicting customers' credit risk (classification task) and real estate price valuation (regression task). The results expose how these metrics can provide a more comprehensive understanding of model predictions and facilitate better communication between decision-makers and stakeholders, thereby increasing the overall transparency and accountability of AI systems.
翻译:随着人工智能技术的飞速发展,治理与监管领域涌现出诸多新挑战。人工智能系统正被整合至各行各业,决策者迫切需要全面且细致地理解这些系统的能力与局限。其中关键环节在于能够解释机器学习模型的结果——这不仅对促进人工智能系统的透明度与可信度至关重要,更是保障机器学习模型合乎伦理训练的基础。本文提出了用于解释分类器与回归器模型预测的新型量化度量框架。所提出的度量标准具有模型无关性,其设计目标在于量化:i. 基于全局与局部特征重要性分布的可解释性因子;ii. 特征对模型输出影响的变异性;iii. 模型决策中特征交互的复杂度。我们采用公开数据集,将所提度量应用于多种机器学习模型,分别针对客户信用风险预测(分类任务)与房地产价格估值(回归任务)。实验结果表明,这些度量能够提供对模型预测更全面的理解,促进决策者与利益相关方之间的有效沟通,从而提升人工智能系统的整体透明度与问责性。