Artificial intelligence (AI) systems utilizing deep neural networks (DNNs) and machine learning (ML) algorithms are widely used for solving important problems in bioinformatics, biomedical informatics, and precision medicine. However, complex DNNs or ML models, which are often perceived as opaque and black-box, can make it difficult to understand the reasoning behind their decisions. This lack of transparency can be a challenge for both end-users and decision-makers, as well as AI developers. Additionally, in sensitive areas like healthcare, explainability and accountability are not only desirable but also legally required for AI systems that can have a significant impact on human lives. Fairness is another growing concern, as algorithmic decisions should not show bias or discrimination towards certain groups or individuals based on sensitive attributes. Explainable artificial intelligence (XAI) aims to overcome the opaqueness of black-box models and provide transparency in how AI systems make decisions. Interpretable ML models can explain how they make predictions and the factors that influence their outcomes. However, most state-of-the-art interpretable ML methods are domain-agnostic and evolved from fields like computer vision, automated reasoning, or statistics, making direct application to bioinformatics problems challenging without customization and domain-specific adaptation. In this paper, we discuss the importance of explainability in the context of bioinformatics, provide an overview of model-specific and model-agnostic interpretable ML methods and tools, and outline their potential caveats and drawbacks. Besides, we discuss how to customize existing interpretable ML methods for bioinformatics problems. Nevertheless, we demonstrate how XAI methods can improve transparency through case studies in bioimaging, cancer genomics, and text mining.
翻译:人工智能系统利用深度神经网络(DNNs)和机器学习(ML)算法被广泛用于解决生物信息学、生物医学信息学和精准医学中的重要问题。然而,复杂的DNNs或ML模型常被视为不透明的“黑箱”,难以理解其决策背后的推理过程。这种透明度的缺失对最终用户、决策者以及AI开发者均构成挑战。此外,在医疗等敏感领域,可解释性和问责性不仅是可取的,更是法律要求——当AI系统对人类生活产生重大影响时,必须满足这些要求。公平性也是日益突出的问题,算法决策不应基于敏感属性对特定群体或个人表现出偏见或歧视。可解释人工智能旨在克服黑箱模型的不透明性,揭示AI系统的决策过程。可解释的ML模型能够阐明预测机制及影响结果的关键因素。然而,现有最先进的可解释ML方法大多具有领域无关性,源自计算机视觉、自动推理或统计学等领域,若不经定制化和领域适应性调整,直接应用于生物信息学问题存在挑战。本文探讨了可解释性在生物信息学中的重要性,概述了模型特定与模型无关的可解释ML方法及工具,并指出其潜在局限与不足。同时,我们讨论了如何为生物信息学问题定制现有可解释ML方法。最后,通过生物成像、癌症基因组学和文本挖掘等案例研究,展示了XAI方法如何提升模型透明度。