Personal informatics (PI) systems, powered by smartphones and wearables, enable people to lead healthier lifestyles by providing meaningful and actionable insights that break down barriers between users and their health information. Today, such systems are used by billions of users for monitoring not only physical activity and sleep but also vital signs and women's and heart health, among others. %Despite their widespread usage, the processing of particularly sensitive personal data, and their proximity to domains known to be susceptible to bias, such as healthcare, bias in PI has not been investigated systematically. Despite their widespread usage, the processing of sensitive PI data may suffer from biases, which may entail practical and ethical implications. In this work, we present the first comprehensive empirical and analytical study of bias in PI systems, including biases in raw data and in the entire machine learning life cycle. We use the most detailed framework to date for exploring the different sources of bias and find that biases exist both in the data generation and the model learning and implementation streams. According to our results, the most affected minority groups are users with health issues, such as diabetes, joint issues, and hypertension, and female users, whose data biases are propagated or even amplified by learning models, while intersectional biases can also be observed.
翻译:个人信息系统(PI)借助智能手机和可穿戴设备,通过提供有意义且可操作的见解,打破用户与其健康信息之间的壁垒,帮助人们养成更健康的生活方式。如今,此类系统已被数十亿用户用于监测身体活动、睡眠、生命体征、女性健康及心脏健康等多方面。尽管其应用广泛,且涉及高度敏感的个人数据处理,并接近医疗等易受偏差影响的领域,但PI中的偏差尚未得到系统研究。尽管PI系统被广泛使用,其对敏感个人数据的处理可能存在偏差,从而引发实践与伦理层面的影响。本文首次对PI系统中的偏差进行全面的实证与分析研究,涵盖原始数据偏差及机器学习全生命周期的偏差。我们采用迄今最详细的框架来探索偏差的不同来源,发现偏差既存在于数据生成环节,也存在于模型学习与实现流程中。研究结果表明,受影响最严重的少数群体是患有健康问题的用户(如糖尿病、关节问题和高血压)以及女性用户,其数据偏差会被学习模型传播甚至放大,同时还可观察到交叉性偏差。