Large language models (LLMs) are becoming bigger to boost performance. However, little is known about how explainability is affected by this trend. This work explores LIME explanations for DeBERTaV3 models of four different sizes on natural language inference (NLI) and zero-shot classification (ZSC) tasks. We evaluate the explanations based on their faithfulness to the models' internal decision processes and their plausibility, i.e. their agreement with human explanations. The key finding is that increased model size does not correlate with plausibility despite improved model performance, suggesting a misalignment between the LIME explanations and the models' internal processes as model size increases. Our results further suggest limitations regarding faithfulness metrics in NLI contexts.
翻译:大型语言模型(LLM)正变得越来越大以提升性能,然而,关于这一趋势对可解释性产生何种影响目前尚不明确。本研究探索了四种不同规模的DeBERTaV3模型在自然语言推理(NLI)和零样本分类(ZSC)任务上的LIME解释。我们根据这些解释对模型内部决策过程的忠实度及其合理性(即与人类解释的一致性)进行评估。关键发现是:尽管模型性能有所提升,但模型规模的增大与合理性之间并无相关性,这表明随着模型规模增大,LIME解释与模型内部过程之间存在偏差。我们的结果进一步揭示了NLI场景下忠实度指标的局限性。