This study delves into the pervasive issue of gender issues in artificial intelligence (AI), specifically within automatic scoring systems for student-written responses. The primary objective is to investigate the presence of gender biases, disparities, and fairness in generally targeted training samples with mixed-gender datasets in AI scoring outcomes. Utilizing a fine-tuned version of BERT and GPT-3.5, this research analyzes more than 1000 human-graded student responses from male and female participants across six assessment items. The study employs three distinct techniques for bias analysis: Scoring accuracy difference to evaluate bias, mean score gaps by gender (MSG) to evaluate disparity, and Equalized Odds (EO) to evaluate fairness. The results indicate that scoring accuracy for mixed-trained models shows an insignificant difference from either male- or female-trained models, suggesting no significant scoring bias. Consistently with both BERT and GPT-3.5, we found that mixed-trained models generated fewer MSG and non-disparate predictions compared to humans. In contrast, compared to humans, gender-specifically trained models yielded larger MSG, indicating that unbalanced training data may create algorithmic models to enlarge gender disparities. The EO analysis suggests that mixed-trained models generated more fairness outcomes compared with gender-specifically trained models. Collectively, the findings suggest that gender-unbalanced data do not necessarily generate scoring bias but can enlarge gender disparities and reduce scoring fairness.
翻译:本研究深入探讨了人工智能(AI)中普遍存在的性别问题,特别聚焦于学生书面回答的自动评分系统。主要目标是调查在混合性别数据集的通用目标训练样本中,AI评分结果是否存在性别偏见、差异及公平性问题。通过使用微调版本的BERT和GPT-3.5,本研究分析了来自六个评估项目中超过1000份由人工评分的男女学生回答。研究采用了三种不同的偏见分析技术:评分准确度差异用于评估偏见,按性别划分的平均分差(MSG)用于评估差异,以及均等几率(EO)用于评估公平性。结果表明,混合训练模型的评分准确度与仅用男性或仅用女性训练模型相比无显著差异,表明不存在显著的评分偏见。与BERT和GPT-3.5的结果一致,我们发现混合训练模型产生的MSG和差异化预测均少于人类评分者。相比之下,与人类评分者相比,按性别特定训练的模型产生了更大的MSG,表明不平衡的训练数据可能导致算法模型扩大性别差异。EO分析表明,混合训练模型比按性别特定训练的模型产生了更公平的结果。综合来看,这些发现表明,性别不平衡的数据不一定产生评分偏见,但可能扩大性别差异并降低评分公平性。