Recent advances and applications of language technology and artificial intelligence have enabled much success across multiple domains like law, medical and mental health. AI-based Language Models, like Judgement Prediction, have recently been proposed for the legal sector. However, these models are strife with encoded social biases picked up from the training data. While bias and fairness have been studied across NLP, most studies primarily locate themselves within a Western context. In this work, we present an initial investigation of fairness from the Indian perspective in the legal domain. We highlight the propagation of learnt algorithmic biases in the bail prediction task for models trained on Hindi legal documents. We evaluate the fairness gap using demographic parity and show that a decision tree model trained for the bail prediction task has an overall fairness disparity of 0.237 between input features associated with Hindus and Muslims. Additionally, we highlight the need for further research and studies in the avenues of fairness/bias in applying AI in the legal sector with a specific focus on the Indian context.
翻译:近期语言技术与人工智能的进步与应用已在法律、医学和心理健康等多个领域取得显著成功。基于AI的语言模型,如判决预测模型,近期被提出用于法律领域。然而,这些模型不可避免地携带了从训练数据中习得的社会偏见。尽管NLP领域已对偏见与公平性展开研究,但大多数研究主要立足于西方语境。本工作首次从印度视角对法律领域的公平性进行初步探究。我们揭示了基于印地语法律文书训练的保释预测模型中算法偏见的传播路径。通过人口统计均等性指标评估公平性差距,结果表明:针对保释预测任务训练的决策树模型中,与印度教徒和穆斯林相关的输入特征之间整体公平性差异达0.237。此外,我们强调需在印度语境下进一步开展法律领域AI应用中公平性/偏见相关问题的研究。