Recent developments in the methods of explainable AI (XAI) allow researchers to explore the inner workings of deep neural networks (DNNs), revealing crucial information about input-output relationships and realizing how data connects with machine learning models. In this paper we explore interpretability of DNN models designed to identify jets coming from top quark decay in high energy proton-proton collisions at the Large Hadron Collider (LHC). We review a subset of existing top tagger models and explore different quantitative methods to identify which features play the most important roles in identifying the top jets. We also investigate how and why feature importance varies across different XAI metrics, how correlations among features impact their explainability, and how latent space representations encode information as well as correlate with physically meaningful quantities. Our studies uncover some major pitfalls of existing XAI methods and illustrate how they can be overcome to obtain consistent and meaningful interpretation of these models. We additionally illustrate the activity of hidden layers as Neural Activation Pattern (NAP) diagrams and demonstrate how they can be used to understand how DNNs relay information across the layers and how this understanding can help to make such models significantly simpler by allowing effective model reoptimization and hyperparameter tuning. These studies not only facilitate a methodological approach to interpreting models but also unveil new insights about what these models learn. Incorporating these observations into augmented model design, we propose the Particle Flow Interaction Network (PFIN) model and demonstrate how interpretability-inspired model augmentation can improve top tagging performance.
翻译:可解释人工智能(XAI)方法的最新进展使研究人员能够深入探索深度神经网络(DNN)的内部运作,揭示输入-输出关系的关键信息,并理解数据如何与机器学习模型相关联。本文探讨了用于识别大型强子对撞机(LHC)高能质子-质子对撞中顶夸克衰变喷注的DNN模型的可解释性。我们回顾了部分现有顶标记模型,并探索了不同的定量方法,以确定哪些特征在识别顶喷注中发挥最关键作用。我们还研究了特征重要性如何及为何在不同XAI指标间存在差异、特征之间的相关性如何影响其可解释性、以及潜在空间表示如何编码信息并与物理上有意义的量相关联。我们的研究揭示了现有XAI方法的一些主要缺陷,并说明了如何克服这些缺陷以获得对这些模型一致且有意义的解释。此外,我们以神经激活模式(NAP)图的形式展示了隐藏层的活动,并演示了如何利用这些图理解DNN如何在层间传递信息,以及这种理解如何通过有效的模型重新优化和超参数调整,使此类模型显著简化。这些研究不仅为解释模型提供了方法论路径,还揭示了关于这些模型学习内容的新见解。通过将这些观察融入增强模型设计中,我们提出了粒子流交互网络(PFIN)模型,并展示了受可解释性启发的模型增强如何提升顶标记性能。