Many evaluation measures are used to evaluate social biases in masked language models (MLMs). However, we find that these previously proposed evaluation measures are lacking robustness in scenarios with limited datasets. This is because these measures are obtained by comparing the pseudo-log-likelihood (PLL) scores of the stereotypical and anti-stereotypical samples using an indicator function. The disadvantage is the limited mining of the PLL score sets without capturing its distributional information. In this paper, we represent a PLL score set as a Gaussian distribution and use Kullback Leibler (KL) divergence and Jensen Shannon (JS) divergence to construct evaluation measures for the distributions of stereotypical and anti-stereotypical PLL scores. Experimental results on the publicly available datasets StereoSet (SS) and CrowS-Pairs (CP) show that our proposed measures are significantly more robust and interpretable than those proposed previously.
翻译:目前,许多评价指标被用于评估掩码语言模型(MLMs)中的社会偏见。然而,我们发现先前提出的这些评价指标在数据集有限的情况下缺乏鲁棒性。这是因为这些指标通过指示函数比较刻板印象样本与反刻板印象样本的伪对数似然(PLL)分数而得到。其缺陷在于对PLL分数集合的挖掘有限,未能捕获其分布信息。本文提出将PLL分数集合表示为高斯分布,并利用Kullback Leibler(KL)散度和Jensen Shannon(JS)散度构建刻板印象与反刻板印象PLL分数分布的评价指标。在公开数据集StereoSet(SS)和CrowS-Pairs(CP)上的实验结果表明,我们提出的指标相较于先前指标具有显著更强的鲁棒性与可解释性。