Masking-based post-hoc explanation methods, such as KernelSHAP and LIME, estimate local feature importance by querying a black-box model under randomized perturbations. This paper formulates this procedure as communication over a query channel, where the latent explanation acts as a message and each masked evaluation is a channel use. Within this framework, the complexity of the explanation is captured by the entropy of the hypothesis class, while the query interface supplies information at a rate determined by an identification capacity per query. We derive a strong converse showing that, if the explanation rate exceeds this capacity, the probability of exact recovery necessarily converges to one in error for any sequence of explainers and decoders. We also prove an achievability result establishing that a sparse maximum-likelihood decoder attains reliable recovery when the rate lies below capacity. A Monte Carlo estimator of mutual information yields a non-asymptotic query benchmark that we use to compare optimal decoding with Lasso- and OLS-based procedures that mirror LIME and KernelSHAP. Experiments reveal a range of query budgets where information theory permits reliable explanations but standard convex surrogates still fail. Finally, we interpret super-pixel resolution and tokenization for neural language models as a source-coding choice that sets the entropy of the explanation and show how Gaussian noise and nonlinear curvature degrade the query channel, induce waterfall and error-floor behavior, and render high-resolution explanations unattainable.


翻译:基于遮掩的事后解释方法(如KernelSHAP和LIME)通过随机扰动下查询黑箱模型来估计局部特征重要性。本文将这一过程建模为在查询信道上进行通信,其中潜在解释充当消息,每次遮掩评估对应一次信道使用。在该框架下,解释的复杂度由假设类的熵捕获,而查询接口以每次查询的识别容量决定的信息速率提供信息。我们推导出一个强逆定理:若解释速率超过该容量,则对于任何解释器与解码器序列,精确恢复的概率必然收敛于误差中的1。同时我们证明了一个可达性结果:当速率低于容量时,稀疏最大似然解码器可实现可靠恢复。互信息的蒙特卡洛估计器提供了非渐近的查询基准,我们将其用于对比最优解码与模拟LIME和KernelSHAP的基于Lasso和OLS的方法。实验揭示了一类查询预算区间,在该区间内信息论允许可靠解释,但标准凸替代方法仍失效。最后,我们将神经语言模型的超像素分辨率和分词解释为设置解释熵的信源编码选择,并展示了高斯噪声与非线性曲率如何劣化查询信道、引发瀑布效应和错误平层现象,并使得高分辨率解释变得不可实现。

0
下载
关闭预览

相关内容

可解释人工智能的基础
专知会员服务
33+阅读 · 2025年10月26日
ISWC2020最佳论文《可解释假信息检测的链接可信度评价》
机器学习的可解释性
专知会员服务
181+阅读 · 2020年8月27日
从信息论的角度来理解损失函数
深度学习每日摘要
17+阅读 · 2019年4月7日
【学界】机器学习模型的“可解释性”到底有多重要?
GAN生成式对抗网络
12+阅读 · 2018年3月3日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
9+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
9+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
14+阅读 · 7月31日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员