It is quite popular nowadays for researchers and data analysts holding different datasets to seek assistance from each other to enhance their modeling performance. We consider a scenario where different learners hold datasets with potentially distinct variables, and their observations can be aligned by a nonprivate identifier. Their collaboration faces the following difficulties: First, learners may need to keep data values or even variable names undisclosed due to, e.g., commercial interest or privacy regulations; second, there are restrictions on the number of transmission rounds between them due to e.g., communication costs. To address these challenges, we develop a two-stage assisted learning architecture for an agent, Alice, to seek assistance from another agent, Bob. In the first stage, we propose a privacy-aware hypothesis testing-based screening method for Alice to decide on the usefulness of the data from Bob, in a way that only requires Bob to transmit sketchy data. Once Alice recognizes Bob's usefulness, Alice and Bob move to the second stage, where they jointly apply a synergistic iterative model training procedure. With limited transmissions of summary statistics, we show that Alice can achieve the oracle performance as if the training were from centralized data, both theoretically and numerically.
翻译:如今,拥有不同数据集的研究人员和数据分析师相互寻求帮助以提升建模性能的做法已相当普遍。我们考虑一个场景:不同学习者持有的数据集包含可能不同的变量,且其观测数据可通过非隐私标识符对齐。其合作面临以下困难:首先,由于商业利益或隐私法规等原因,学习者可能需要隐瞒数据值甚至变量名称;其次,受通信成本等因素限制,双方之间的传输轮次存在约束。为解决这些挑战,我们为代理Alice开发了一种两阶段辅助学习架构,使其可从另一代理Bob处寻求帮助。第一阶段,我们提出一种基于隐私感知假设检验的筛选方法,使Alice仅需Bob传输概略数据即可判断其数据的有效性。一旦Alice确认Bob的数据有用,双方进入第二阶段,在此阶段联合应用协同迭代模型训练流程。通过有限次传输汇总统计量,我们从理论和数值上证明,Alice能够达到仿佛数据来自集中训练般的预言性能。