Dyadic regression models, which predict real-valued outcomes for pairs of entities, are fundamental in many domains (e.g. predicting the rating of a user to a product in Recommender Systems) and promising and under exploration in many others (e.g. approximating the adequate dosage of a drug for a patient in personalized pharmacology). In this work, we demonstrate that non-uniformity in the observed value distributions of individual entities leads to severely biased predictions in state-of-the-art models, skewing predictions towards the average of observed past values for the entity and providing worse-than-random predictive power in eccentric yet equally important cases. We show that the usage of global error metrics like Root Mean Squared Error (RMSE) and Mean Absolute Error (MAE) is insufficient to capture this phenomenon, which we name eccentricity bias, and we introduce Eccentricity-Area Under the Curve (EAUC) as a new complementary metric that can quantify it in all studied models and datasets. We also prove the adequateness of EAUC by using naive de-biasing corrections to demonstrate that a lower model bias correlates with a lower EAUC and vice-versa. This work contributes a bias-aware evaluation of dyadic regression models to avoid potential unfairness and risks in critical real-world applications of such systems.
翻译:二元回归模型用于预测实体对之间连续值结果,是多个领域的基础(例如在推荐系统中预测用户对产品的评分),并在其他领域展现出前景且尚在探索中(如个性化药理学中预测患者合适的药物剂量)。本研究表明,单个实体观测值分布的非均匀性会导致最先进模型产生严重偏差,使预测结果偏向于该实体历史观测值的平均值,并在偏离常规但同样重要的案例中提供弱于随机猜测的预测能力。我们证明,使用全局误差指标(如均方根误差RMSE和平均绝对误差MAE)不足以捕捉这一现象——我们将其命名为“离群偏差”,并引入“离群曲线下面积”(EAUC)作为新的补充性指标,以量化所有研究模型与数据集中的此类偏差。此外,我们通过简单的去偏校正方法验证了EAUC的适用性:模型偏差降低与EAUC下降呈正相关,反之亦然。本研究为二元回归模型提供了偏差感知评估方法,以避免此类系统在关键现实应用中可能产生的不公平性与风险。