In this work, we propose a novel tree-based explanation technique, PEACH (Pretrained-embedding Explanation Across Contextual and Hierarchical Structure), that can explain how text-based documents are classified by using any pretrained contextual embeddings in a tree-based human-interpretable manner. Note that PEACH can adopt any contextual embeddings of the PLMs as a training input for the decision tree. Using the proposed PEACH, we perform a comprehensive analysis of several contextual embeddings on nine different NLP text classification benchmarks. This analysis demonstrates the flexibility of the model by applying several PLM contextual embeddings, its attribute selections, scaling, and clustering methods. Furthermore, we show the utility of explanations by visualising the feature selection and important trend of text classification via human-interpretable word-cloud-based trees, which clearly identify model mistakes and assist in dataset debugging. Besides interpretability, PEACH outperforms or is similar to those from pretrained models.
翻译:本文提出了一种新颖的树形解释技术——PEACH(基于预训练嵌入的上下文与层次结构解释方法),该技术能够以基于树的人类可解释方式,利用任意预训练上下文嵌入来解释文本型文档的分类过程。值得注意的是,PEACH可将预训练语言模型(PLM)的任何上下文嵌入作为决策树的训练输入。通过所提出的PEACH方法,我们对九种不同自然语言处理文本分类基准任务上的多个上下文嵌入进行了全面分析。该分析通过应用多种PLM上下文嵌入及其属性选择、缩放和聚类方法,展现了模型的灵活性。此外,我们通过基于人类可解释词云树的特征可视化与重要趋势展示,揭示了模型解释的实用性,从而清晰识别模型错误并辅助数据集调试。除可解释性外,PEACH的性能优于或等同于预训练模型。