Interpretability methods in NLP aim to provide insights into the semantics underlying specific system architectures. Focusing on word embeddings, we present a supervised-learning method that, for a given domain (e.g., sports, professions), identifies a subset of model features that strongly improve prediction of human similarity judgments. We show this method keeps only 20-40% of the original embeddings, for 8 independent semantic domains, and that it retains different feature sets across domains. We then present two approaches for interpreting the semantics of the retained features. The first obtains the scores of the domain words (co-hyponyms) on the first principal component of the retained embeddings, and extracts terms whose co-occurrence with the co-hyponyms tracks these scores' profile. This analysis reveals that humans differentiate e.g. sports based on how gender-inclusive and international they are. The second approach uses the retained sets as variables in a probing task that predicts values along 65 semantically annotated dimensions for a dataset of 535 words. The features retained for professions are best at predicting cognitive, emotional and social dimensions, whereas features retained for fruits or vegetables best predict the gustation (taste) dimension. We discuss implications for alignment between AI systems and human knowledge.
翻译:自然语言处理中的可解释性方法旨在揭示特定系统架构背后的语义内涵。针对词嵌入,我们提出一种监督学习方法:在给定领域(如体育、职业)中,该方法能识别出显著提升人类相似性判断预测效果的模型特征子集。实验表明,在8个独立语义领域内,该方法仅保留原始嵌入的20-40%,且不同领域保留的特征集具有差异性。我们进而提出两种解释保留特征语义的方法:第一种方法获取领域词(共下义词)在保留嵌入第一主成分上的得分,并提取与共下义词共现模式追踪这些得分分布特征的术语。该分析揭示人类区分体育项目时主要依据其性别包容性与国际性程度。第二种方法将保留特征集作为探测任务的变量,对535个词语数据集进行65个语义标注维度的值预测。其中,职业领域保留的特征在认知、情感和社会维度上的预测表现最优,而水果或蔬菜领域保留的特征则最佳地预测了味觉维度。我们探讨了该研究对人工智能系统与人类知识对齐的启示。