Machine learning methods for estimating heterogeneous treatment effects (HTE) facilitate large-scale personalized decision-making across various domains such as healthcare, policy making, education, and more. Current machine learning approaches for HTE require access to substantial amounts of data per treatment, and the high costs associated with interventions makes centrally collecting so much data for each intervention a formidable challenge. To overcome this obstacle, in this work, we propose a novel framework for collaborative learning of HTE estimators across institutions via Federated Learning. We show that even under a diversity of interventions and subject populations across clients, one can jointly learn a common feature representation, while concurrently and privately learning the specific predictive functions for outcomes under distinct interventions across institutions. Our framework and the associated algorithm are based on this insight, and leverage tabular transformers to map multiple input data to feature representations which are then used for outcome prediction via multi-task learning. We also propose a novel way of federated training of personalised transformers that can work with heterogeneous input feature spaces. Experimental results on real-world clinical trial data demonstrate the effectiveness of our method.
翻译:用于估计异质性处理效应的机器学习方法可促进医疗、政策制定、教育等多个领域的大规模个性化决策。当前针对HTE的机器学习方法需要获取每个处理组的大量数据,而干预措施的高昂成本使得集中收集每种干预的如此大量数据成为一项艰巨挑战。为克服这一障碍,我们提出了一种基于联邦学习的跨机构HTE估计器协作学习新框架。研究表明,即使各客户端面临的干预措施与受试人群存在多样性,仍可联合学习公共特征表示,同时以隐私保护方式分别学习各机构不同干预措施下结果变量的特定预测函数。该框架及其相关算法基于这一洞见,利用表格型Transformer将多源输入数据映射为特征表示,进而通过多任务学习进行结果预测。我们还提出了一种新型的个性化Transformer联邦训练方法,可处理异构输入特征空间。基于真实临床试验数据的实验结果验证了本方法的有效性。