Deep Learning has achieved tremendous results by pushing the frontier of automation in diverse domains. Unfortunately, current neural network architectures are not explainable by design. In this paper, we propose a novel method that trains deep hypernetworks to generate explainable linear models. Our models retain the accuracy of black-box deep networks while offering free lunch explainability by design. Specifically, our explainable approach requires the same runtime and memory resources as black-box deep models, ensuring practical feasibility. Through extensive experiments, we demonstrate that our explainable deep networks are as accurate as state-of-the-art classifiers on tabular data. On the other hand, we showcase the interpretability of our method on a recent benchmark by empirically comparing prediction explainers. The experimental results reveal that our models are not only as accurate as their black-box deep-learning counterparts but also as interpretable as state-of-the-art explanation techniques.
翻译:深度学习通过在多个领域推动自动化前沿取得了巨大成就。然而,当前神经网络架构并非设计为可解释的。本文提出了一种创新方法,即训练深度超网络生成可解释的线性模型。我们的模型在保留黑盒深度网络精度的同时,通过设计实现免费的可解释性。具体而言,该可解释方法在运行时和内存资源消耗上与黑盒深度模型相当,确保了实际可行性。通过大量实验,我们证明该可解释深度网络在表格数据上的准确率与最先进的分类器持平。另一方面,我们通过经验比较预测解释器,在近期基准测试中展示了方法的可解释性。实验结果表明,我们的模型不仅与黑盒深度学习模型同样准确,而且其可解释性可媲美最先进的解释技术。