Machine learning models are often personalized with information that is protected, sensitive, self-reported, or costly to acquire. These models use information about people but do not facilitate nor inform their consent. Individuals cannot opt out of reporting personal information to a model, nor tell if they benefit from personalization in the first place. We introduce a family of classification models, called participatory systems, that let individuals opt into personalization at prediction time. We present a model-agnostic algorithm to learn participatory systems for personalization with categorical group attributes. We conduct a comprehensive empirical study of participatory systems in clinical prediction tasks, benchmarking them with common approaches for personalization and imputation. Our results demonstrate that participatory systems can facilitate and inform consent while improving performance and data use across all groups who report personal data.
翻译:机器学习模型通常利用受保护、敏感、自我报告或获取成本高昂的信息进行个性化。这些模型虽使用个人信息,但既不促进也不告知用户的知情同意。个人无法选择不向模型报告个人信息,也无法判断自己是否从个性化中受益。我们提出一类称为参与式系统的分类模型,允许个体在预测时自主选择是否参与个性化。我们提出一种与模型无关的算法,用于学习具有分类群体属性的个性化参与式系统。我们在临床预测任务中对参与式系统进行了全面的实证研究,将其与常见的个性化和插补方法进行基准对比。结果表明,参与式系统能够促进并告知用户知情同意,同时在所有报告个人数据的群体中提升模型性能和数据利用效率。