To address the interpretability challenge in machine learning (ML) systems, counterfactual explanations (CEs) have emerged as a promising solution. CEs are unique as they provide workable suggestions to users, in addition to explaining why a certain outcome was predicted. The application of CEs encounters two main challenges: general user preferences and variable ML systems. User preferences tend to be general rather than specific, and CEs need to be adaptable to variable ML models while maintaining robustness even as these models change. Facing these challenges, we present a solution rooted in validated general user preferences, which are derived from thorough user research. We map these preferences to the properties of CEs. Additionally, we introduce a novel method, \uline{T}ree-based \uline{C}onditions \uline{O}ptional \uline{L}inks (T-COL), which incorporates two optional structures and multiple condition groups for generating CEs adaptable to general user preferences. Meanwhile, we employ T-COL to enhance the robustness of CEs with specific conditions, making them more valid even when the ML model is replaced. Our experimental comparisons under different user preferences show that T-COL outperforms all baselines, including Large Language Models which are shown to be able to generate counterfactuals.
翻译:为解决机器学习(ML)系统中的可解释性挑战,反事实解释(CEs)已成为一种有前景的解决方案。CEs的独特之处在于,除了解释为何预测出某种结果外,还能为用户提供可操作的改进建议。CEs的应用面临两大挑战:一般用户偏好与可变ML系统。用户偏好往往具有一般性而非特异性,且CEs需要适应可变的ML模型,即使这些模型发生变更也能保持稳健性。针对这些挑战,我们提出一种基于经过验证的一般用户偏好的解决方案——这些偏好源自深入的用户研究,并将这些偏好映射到CEs的属性上。此外,我们引入一种创新方法\uline{T}ree-based \uline{C}onditions \uline{O}ptional \uline{L}inks(T-COL),该方法包含两种可选结构和多个条件组,用于生成适应一般用户偏好的CEs。同时,我们利用T-COL增强具有特定条件的CEs的稳健性,使其在ML模型被替换后仍保持有效性。在不同用户偏好下的实验对比表明,T-COL优于所有基线方法,包括已被证明能够生成反事实的大语言模型。