Recent advances and applications of language technology and artificial intelligence have enabled much success across multiple domains like law, medical and mental health. AI-based Language Models, like Judgement Prediction, have recently been proposed for the legal sector. However, these models are strife with encoded social biases picked up from the training data. While bias and fairness have been studied across NLP, most studies primarily locate themselves within a Western context. In this work, we present an initial investigation of fairness from the Indian perspective in the legal domain. We highlight the propagation of learnt algorithmic biases in the bail prediction task for models trained on Hindi legal documents. We evaluate the fairness gap using demographic parity and show that a decision tree model trained for the bail prediction task has an overall fairness disparity of 0.237 between input features associated with Hindus and Muslims. Additionally, we highlight the need for further research and studies in the avenues of fairness/bias in applying AI in the legal sector with a specific focus on the Indian context.
翻译:近期语言技术与人工智能的进步与广泛应用已在法律、医疗、心理健康等多个领域取得显著成功。基于人工智能的语言模型(如判决预测)最近被提出应用于法律领域。然而,这些模型充斥着从训练数据中习得的编码社会偏见。尽管自然语言处理领域已对偏见与公平性展开研究,多数研究主要立足于西方语境。本研究首次从印度视角对法律领域的公平性进行初步探讨。我们揭示了在印地语法律文档上训练的保释预测模型中习得算法偏见的传播路径。通过人口统计均等性评估公平性差距,结果显示,针对保释预测任务训练的决策树模型,在与印度教和穆斯林相关的输入特征之间,整体公平性差距为0.237。此外,我们强调需在人工智能应用于法律领域的公平性/偏见研究方向开展进一步研究,并特别关注印度语境。