Traditional Federated Learning (FL) follows a server-dominated cooperation paradigm which narrows the application scenarios of FL and decreases the enthusiasm of data holders to participate. To fully unleash the potential of FL, we advocate rethinking the design of current FL frameworks and extending it to a more generalized concept: Open Federated Learning Platforms, positioned as a crowdsourcing collaborative machine learning infrastructure for all Internet users. We propose two reciprocal cooperation frameworks to achieve this: query-based FL and contract-based FL. In this survey, we conduct a comprehensive review of the feasibility of constructing open FL platforms from both technical and legal perspectives. We begin by reviewing the definition of FL and summarizing its inherent limitations, including server-client coupling, low model reusability, and non-public. In particular, we introduce a novel taxonomy to streamline the analysis of model license compatibility in FL studies that involve batch model reusing methods, including combination, amalgamation, distillation, and generation. This taxonomy provides a feasible solution for identifying the corresponding licenses clauses and facilitates the analysis of potential legal implications and restrictions when reusing models. Through this survey, we uncover the current dilemmas faced by FL and advocate for the development of sustainable open FL platforms. We aim to provide guidance for establishing such platforms in the future while identifying potential limitations that need to be addressed.
翻译:传统的联邦学习遵循以服务器为主导的合作范式,这缩小了联邦学习的应用场景,并降低了数据持有者的参与积极性。为充分释放联邦学习的潜力,我们倡导重新思考当前联邦学习框架的设计,并将其扩展至一个更广义的概念:面向开放的联邦学习平台,将其定位为面向所有互联网用户的众包协作机器学习基础设施。为实现这一目标,我们提出了两种互惠合作框架:基于查询的联邦学习和基于合约的联邦学习。在本综述中,我们从技术与法律两个视角,对构建开放联邦学习平台的可行性进行了全面审视。我们首先回顾联邦学习的定义,并概括其固有局限性,包括服务器-客户端耦合、模型复用性低以及非公开性。特别是,针对涉及批量模型复用方法(包括组合、融合、蒸馏和生成)的联邦学习研究,我们引入了一种新颖的分类法,以简化模型许可兼容性的分析。该分类法为识别相应许可条款提供了可行方案,并有助于分析复用模型时潜在的法律影响与限制。通过本综述,我们揭示了当前联邦学习所面临的困境,并倡导开发可持续的开放联邦学习平台。我们旨在为未来建立此类平台提供指导,同时识别有待解决的潜在局限性。