Shapley values are a cornerstone of explainable AI, yet their proliferation into competing formulations has created a fragmented landscape with little consensus on practical deployment. While theoretical differences are well-documented, evaluation remains reliant on quantitative proxies whose alignment with human utility is unverified. In this work, we use a unified amortized framework to isolate semantic differences between eight Shapley variants under the low-latency constraints of operational risk workflows. We conduct a large-scale empirical evaluation across four risk datasets and a realistic fraud-detection environment involving professional analysts and 3,735 case reviews. Our results reveal a fundamental misalignment: standard quantitative metrics, such as sparsity and faithfulness, are decoupled from human-perceived clarity and decision utility. Furthermore, while no formulation improved objective analyst performance, explanations consistently increased decision confidence, signaling a critical risk of automation bias in high-stakes settings. These findings suggest that current evaluation proxies are insufficient for predicting downstream human impact, and we provide evidence-based guidance for selecting formulations and metrics in operational decision systems.
翻译:Shapley值是可解释人工智能的基石,但其衍生出的多种竞争性算法导致领域碎片化,对实际部署缺乏共识。尽管理论差异已有充分研究,其评估仍依赖与人类效用一致性未经验证的定量代理指标。本研究采用统一摊销框架,在工作流低延迟约束下隔离八种Shapley变体的语义差异,基于四个风险数据集与含专业分析师及3,735例案例审查的真实欺诈检测环境开展大规模实证评估。结果表明存在根本性错位:稀疏性、忠实度等标准定量指标与人类感知清晰度及决策效用脱节。此外,尽管无算法能提升分析师客观表现,解释却持续增强决策信心,预示高风险场景中存在自动化偏差风险。这些发现表明,现有评估代理指标不足以预测下游影响,我们为实际决策系统中的算法与指标选择提供了循证指南。