Conversion rate (CVR) estimation aims to predict the probability of conversion event after a user has clicked an ad. Typically, online publisher has user browsing interests and click feedbacks, while demand-side advertising platform collects users' post-click behaviors such as dwell time and conversion decisions. To estimate CVR accurately and protect data privacy better, vertical federated learning (vFL) is a natural solution to combine two sides' advantages for training models, without exchanging raw data. Both CVR estimation and applied vFL algorithms have attracted increasing research attentions. However, standardized and systematical evaluations are missing: due to the lack of standardized datasets, existing studies adopt public datasets to simulate a vFL setting via hand-crafted feature partition, which brings challenges to fair comparison. We introduce FedAds, the first benchmark for CVR estimation with vFL, to facilitate standardized and systematical evaluations for vFL algorithms. It contains a large-scale real world dataset collected from Alibaba's advertising platform, as well as systematical evaluations for both effectiveness and privacy aspects of various vFL algorithms. Besides, we also explore to incorporate unaligned data in vFL to improve effectiveness, and develop perturbation operations to protect privacy well. We hope that future research work in vFL and CVR estimation benefits from the FedAds benchmark.
翻译:转化率(CVR)估计旨在预测用户点击广告后发生转化事件的概率。通常,在线发布平台掌握用户的浏览兴趣和点击反馈,而需求方广告平台收集用户点击后的行为数据(如停留时间和转化决策)。为在保护数据隐私的同时准确估计CVR,纵向联邦学习(vFL)成为自然解决方案——它能在不交换原始数据的前提下,融合双方优势进行模型训练。尽管CVR估计与vFL算法均已引发广泛研究关注,但目前仍缺乏标准化、系统化的评估体系:由于缺少标准化数据集,现有研究通常采用公开数据集并手动划分特征来模拟vFL场景,这给公平比较带来了挑战。为此,我们提出FedAds——首个面向纵向联邦学习CVR估计的基准测试,旨在为vFL算法提供标准化、系统化的评估框架。该基准包含从阿里巴巴广告平台收集的大规模真实数据集,并系统评估了多种vFL算法的有效性与隐私保护能力。此外,我们探索了在vFL中整合非对齐数据以提升效果的方法,并开发了扰动操作以强化隐私保护。期望未来vFL与CVR估计领域的研究工作能受益于FedAds基准测试。