Word-level saliency explanations ("heat maps over words") are often used to communicate feature-attribution in text-based models. Recent studies found that superficial factors such as word length can distort human interpretation of the communicated saliency scores. We conduct a user study to investigate how the marking of a word's neighboring words affect the explainee's perception of the word's importance in the context of a saliency explanation. We find that neighboring words have significant effects on the word's importance rating. Concretely, we identify that the influence changes based on neighboring direction (left vs. right) and a-priori linguistic and computational measures of phrases and collocations (vs. unrelated neighboring words). Our results question whether text-based saliency explanations should be continued to be communicated at word level, and inform future research on alternative saliency explanation methods.
翻译:词级显著性解释(“词语热力图”)常用于传达文本模型中的特征归因。近期研究发现,词长等表面因素会扭曲人类对所传达显著性分数的理解。我们开展了一项用户研究,探究词语的邻词标记如何影响解释对象在显著性解释语境中对词语重要性的感知。研究发现邻词对词语重要性评级具有显著影响。具体而言,我们识别出这种影响会因邻词方向(左侧vs右侧)以及短语和搭配的先验语言与计算度量(vs.无关邻词)而变化。我们的研究结果质疑是否应继续以词级形式传达基于文本的显著性解释,并为未来关于替代性显著性解释方法的研究提供参考。