Masking-based post-hoc explanation methods, such as KernelSHAP and LIME, estimate local feature importance by querying a black-box model under randomized perturbations. This paper formulates this procedure as communication over a query channel, where the latent explanation acts as a message and each masked evaluation is a channel use. Within this framework, the complexity of the explanation is captured by the entropy of the hypothesis class, while the query interface supplies information at a rate determined by an identification capacity per query. We derive a strong converse showing that, if the explanation rate exceeds this capacity, the probability of exact recovery necessarily converges to one in error for any sequence of explainers and decoders. We also prove an achievability result establishing that a sparse maximum-likelihood decoder attains reliable recovery when the rate lies below capacity. A Monte Carlo estimator of mutual information yields a non-asymptotic query benchmark that we use to compare optimal decoding with Lasso- and OLS-based procedures that mirror LIME and KernelSHAP. Experiments reveal a range of query budgets where information theory permits reliable explanations but standard convex surrogates still fail. Finally, we interpret super-pixel resolution and tokenization for neural language models as a source-coding choice that sets the entropy of the explanation and show how Gaussian noise and nonlinear curvature degrade the query channel, induce waterfall and error-floor behavior, and render high-resolution explanations unattainable.
翻译:基于遮掩的事后解释方法(如KernelSHAP和LIME)通过随机扰动下查询黑箱模型来估计局部特征重要性。本文将这一过程建模为在查询信道上进行通信,其中潜在解释充当消息,每次遮掩评估对应一次信道使用。在该框架下,解释的复杂度由假设类的熵捕获,而查询接口以每次查询的识别容量决定的信息速率提供信息。我们推导出一个强逆定理:若解释速率超过该容量,则对于任何解释器与解码器序列,精确恢复的概率必然收敛于误差中的1。同时我们证明了一个可达性结果:当速率低于容量时,稀疏最大似然解码器可实现可靠恢复。互信息的蒙特卡洛估计器提供了非渐近的查询基准,我们将其用于对比最优解码与模拟LIME和KernelSHAP的基于Lasso和OLS的方法。实验揭示了一类查询预算区间,在该区间内信息论允许可靠解释,但标准凸替代方法仍失效。最后,我们将神经语言模型的超像素分辨率和分词解释为设置解释熵的信源编码选择,并展示了高斯噪声与非线性曲率如何劣化查询信道、引发瀑布效应和错误平层现象,并使得高分辨率解释变得不可实现。