AI governance frameworks increasingly emphasize fairness, transparency, accountability, and lifecycle risk management in high-stakes domains. However, many current approaches remain observational, relying on static metric reporting, post-hoc auditing, and monitoring dashboards without directly governing deployment readiness, remediation progression, escalation states, or assurance-driven deployment control. This paper introduces Operational AI Deployment Assurance (OADA), a governance framework for translating fairness disagreement, subgroup instability, threshold sensitivity, remediation outcomes, and operational uncertainty into deployment-oriented assurance decisions. Building on prior work on the Fairness Disagreement Index (FDI) and FairRisk-FDI, OADA reframes governance uncertainty as an operational concern within AI deployment pipelines rather than a byproduct of metric disagreement. The framework introduces Deployment Assurance Scores, Deployment Readiness Classifications, Threshold Stability Zones, Governance Escalation States, and remediation-aware assurance progression. These constructs support lifecycle-oriented governance decisions across high-stakes settings by connecting evaluation outputs to deployment-state interpretation, reassessment, escalation, and operational control. Through deployment-oriented evaluation across facial recognition systems, with discussion extended to healthcare AI as a representative high-stakes domain, the paper demonstrates how systems may appear acceptable under isolated fairness or performance metrics while still exhibiting instability that affects deployment readiness. The proposed framework positions operational deployment assurance as a governance layer between evaluation and real-world AI deployment.
翻译:人工智能治理框架日益强调高风险领域中的公平性、透明度、问责制和全生命周期风险管理。然而,当前许多方法仍停留在观察层面,依赖静态指标报告、事后审计和监控仪表盘,未能直接管理部署就绪度、修复进程、升级状态或基于保障的部署控制。本文提出操作化部署保障(OADA),这是一种将公平性分歧、子群不稳定性、阈值敏感性、修复结果及操作不确定性转化为面向部署的保障决策的治理框架。基于先前关于公平性分歧指数(FDI)和FairRisk-FDI的研究,OADA将治理不确定性重新定义为AI部署流水线中的操作性问题,而非指标分歧的副产品。该框架引入了部署保障评分、部署就绪度分类、阈值稳定区间、治理升级状态以及考虑修复的保障进程。这些构造通过将评估输出与部署状态解读、重新评估、升级和操作控制相连接,支持跨高风险场景的生命周期治理决策。通过在面部识别系统中开展面向部署的评估,并将讨论延伸至代表性高风险领域(医疗AI),本文展示了系统在孤立的公平性或性能指标下可能看似可接受,但仍存在影响部署就绪度的不稳定性。所提出的框架将操作化部署保障定位为评估与现实世界AI部署之间的治理层。