In high-speed rail (HSR) systems, federated learning (FL) enables cross-departmental flow prediction without sharing raw data. However, existing schemes suffer from two key limitations: (1) insufficient incentives, leading to free-riding and model poisoning; and (2) centralized aggregation, which introduces a single point of failure. We propose a secure and efficient framework SI-ChainFL that addresses these issues by combining contribution-aware incentives with decentralized aggregation. First, we quantify client contributions using a Shapley value metric that jointly considers rare-event utility, data diversity, data quality, and timeliness. To reduce computational overhead, we further develop a rare positive driven client clustering strategy to accelerate Shapley estimation. Moreover, we design a blockchain-based consensus protocol for decentralized aggregation, where aggregation eligibility is tied to Shapley incentives. This design motivates clients to submit high-quality updates and enables efficient and secure global aggregation. Experiments on MNIST, CIFAR 10 and CIFAR 100, and a HSR flow dataset show that SI ChainFL remains effective under 90% malicious clients in PA attacks, achieving 14.12% higher accuracy than RAGA. Theoretical analysis further guarantees an upper bound on performance
翻译:在高速铁路(HSR)系统中,联邦学习(FL)能够在不共享原始数据的情况下实现跨部门流量预测。然而,现有方案存在两个关键局限:(1)激励不足,导致搭便车和模型投毒;(2)集中式聚合,引入了单点故障。我们提出了一个安全高效的框架SI-ChainFL,通过将贡献感知激励与去中心化聚合相结合来解决这些问题。首先,我们使用Shapley值度量来量化客户端贡献,该度量综合考虑了稀有事件效用、数据多样性、数据质量和时效性。为降低计算开销,我们进一步提出了一种稀有正例驱动的客户端聚类策略以加速Shapley值估计。此外,我们设计了一种基于区块链的共识协议用于去中心化聚合,其中聚合资格与Shapley激励挂钩。该设计激励客户端提交高质量更新,并实现高效安全的全局聚合。在MNIST、CIFAR 10、CIFAR 100以及一个HSR流量数据集上的实验表明,SI-ChainFL在PA攻击下面对90%恶意客户端时仍保持有效,其准确率比RAGA高出14.12%。理论分析进一步保证了性能上界。