Content Warning: This work contains examples that potentially implicate stereotypes, associations, and other harms that could be offensive to individuals in certain social groups.} Large pre-trained language models are acknowledged to carry social biases towards different demographics, which can further amplify existing stereotypes in our society and cause even more harm. Text-to-SQL is an important task, models of which are mainly adopted by administrative industries, where unfair decisions may lead to catastrophic consequences. However, existing Text-to-SQL models are trained on clean, neutral datasets, such as Spider and WikiSQL. This, to some extent, cover up social bias in models under ideal conditions, which nevertheless may emerge in real application scenarios. In this work, we aim to uncover and categorize social biases in Text-to-SQL models. We summarize the categories of social biases that may occur in structured data for Text-to-SQL models. We build test benchmarks and reveal that models with similar task accuracy can contain social biases at very different rates. We show how to take advantage of our methodology to uncover and assess social biases in the downstream Text-to-SQL task. We will release our code and data.
翻译:内容警告:本作品包含可能涉及刻板印象、联想及其他对特定社会群体成员具有冒犯性的伤害性示例。大型预训练语言模型被公认为对不同人群存在社会偏见,这种偏见可能进一步放大我们现存的刻板印象并造成更多伤害。Text-to-SQL是一项重要任务,其模型主要被行政行业采用,在这些行业中不公正的决策可能导致灾难性后果。然而,现有Text-to-SQL模型均基于如Spider和WikiSQL等干净、中立的数据集进行训练,这在一定程度上掩盖了模型在理想条件下的社会偏见,而这些偏见在实际应用场景中仍可能显现。本研究旨在揭示与分类Text-to-SQL模型中的社会偏见。我们总结了Text-to-SQL模型在结构化数据中可能出现的偏见类别,构建了测试基准,并发现任务精度相近的模型其社会偏见比例可能存在显著差异。我们展示了如何利用所提出的方法论来揭示与评估下游Text-to-SQL任务中的社会偏见。我们将公开发布代码与数据集。