Fairness is one of the socio-technical concerns that must be addressed in AI-based systems. Unfair AI-based systems, particularly unfair AI-based mobile apps, can pose difficulties for a significant proportion of the global population. This paper aims to analyze fairness concerns in AI-based app reviews. We first manually constructed a ground-truth dataset, including 1,132 fairness and 1,473 non-fairness reviews. Leveraging the ground-truth dataset, we developed and evaluated a set of machine learning and deep learning models that distinguish fairness reviews from non-fairness reviews. Our experiments show that our best-performing model can detect fairness reviews with a precision of 94%. We then applied the best-performing model on approximately 9.5M reviews collected from 108 AI-based apps and identified around 92K fairness reviews. Next, applying the K-means clustering technique to the 92K fairness reviews, followed by manual analysis, led to the identification of six distinct types of fairness concerns (e.g., 'receiving different quality of features and services in different platforms and devices' and 'lack of transparency and fairness in dealing with user-generated content'). Finally, the manual analysis of 2,248 app owners' responses to the fairness reviews identified six root causes (e.g., 'copyright issues') that app owners report to justify fairness concerns.
翻译:公平性是AI系统中必须解决的社会技术关切之一。不公平的AI系统,尤其是不公平的基于AI的移动应用,可能给全球相当大比例的人口带来困扰。本文旨在分析基于AI的应用评论中的公平性问题。我们首先手动构建了一个包含1,132条公平性评论和1,473条非公平性评论的基准数据集。利用该数据集,我们开发并评估了一系列能够区分公平性与非公平性评论的机器学习和深度学习模型。实验表明,我们性能最佳的模型能以94%的精确度检测公平性评论。随后,我们将该模型应用于从108个基于AI的应用中收集的大约950万条评论,识别出约9.2万条公平性评论。接着,通过对这9.2万条公平性评论应用K-means聚类技术并进行人工分析,我们识别出六种不同类型的公平性问题(例如“在不同平台和设备上接收不同质量的功能和服务”以及“在处理用户生成内容时缺乏透明度和公平性”)。最后,通过对2,248条应用开发者对公平性评论的回复进行人工分析,我们确定了应用开发者用来解释公平性问题的六种根本原因(例如“版权问题”)。