Statistical decision algorithms are increasingly deployed in domains where ground-truth labels are hard to obtain, such as hiring, university admissions, and content moderation. In these settings, models are typically trained on historical human evaluations -- for example, using past hiring decisions as a proxy for true applicant quality. However, if past evaluations unjustly favor certain groups, models trained on these labels may inherit those biases. To address this problem, we propose basing predictions on rubric embeddings, a representation framework that replaces standard black-box embeddings with features derived from expert-defined criteria that align with the underlying construct of interest. By anchoring predictions to semantically meaningful dimensions, this approach guards against biased proxy signals. We provide both theoretical and empirical evidence that rubric embeddings mitigate label bias under plausible conditions. Empirically, we evaluate our method on a novel dataset of applications to a large master's program. We find that models trained on rubric embeddings reduce group disparities while improving measures of cohort quality. Our results suggest that basing predictions on interpretable, domain-grounded representations offers a practical approach to learning in the presence of biased labels.
翻译:统计决策算法越来越多地部署于难以获得真实标签的领域,例如招聘、大学录取和内容审核。在这些场景中,模型通常基于历史人工评估进行训练——例如,以过去的招聘决策作为申请人真实质量的代理指标。然而,若历史评估不公正地偏向特定群体,基于这些标签训练的模型可能继承此类偏差。为解决这一问题,我们提出基于量规嵌入(rubric embeddings)进行预测——该表示框架用专家定义、与底层目标构念一致的标准特征替代标准黑盒嵌入。通过将预测锚定于语义上有意义的维度,该方法可防范有偏差的代理信号。我们从理论和实证两方面证明,量规嵌入在合理条件下能缓解标签偏差。实证方面,我们在某大型硕士项目申请数据集上评估了该方法。研究发现,基于量规嵌入训练的模型在减少群体差异的同时提升了群体质量指标。结果表明,将预测基于可解释、领域扎根的表示,为在存在偏差标签的情况下进行学习提供了实用方案。