Recent years have witnessed a surge in deep learning research, marked by the introduction of expansive generative models like OpenAI's SORA and GPT, Meta AI's LLAMA series, and Google's FLAN, BART, and Gemini models. However, the rapid advancement of large models (LM) has intensified the demand for computing resources, particularly GPUs, which are crucial for their parallel processing capabilities. This demand is exacerbated by limited GPU availability due to supply chain delays and monopolistic acquisition by major tech firms. Distributed Machine Learning (DML) methods, such as Federated Learning (FL), mitigate these challenges by partitioning data and models across multiple servers, though implementing optimizations like tensor and pipeline parallelism remains complex. Blockchain technology emerges as a promising solution, ensuring data integrity, scalability, and trust in distributed computing environments, but still lacks guidance on building practical DML systems. In this paper, we propose a \textit{trustworthy distributed machine learning} (TDML) framework that leverages blockchain to coordinate remote trainers and validate workloads, achieving privacy, transparency, and efficient model training across public remote computing resources. Experimental validation demonstrates TDML's efficacy in overcoming performance limitations and malicious node detection, positioning it as a robust solution for scalable and secure distributed machine learning.
翻译:近年来,深度学习研究蓬勃发展,以OpenAI的SORA和GPT、Meta AI的LLAMA系列以及Google的FLAN、BART和Gemini等大规模生成模型的推出为标志。然而,大模型(LM)的快速发展加剧了对计算资源(尤其是GPU)的需求,GPU因其并行处理能力而至关重要。由于供应链延迟和大型科技公司的垄断性采购,GPU供应有限,进一步加剧了这一需求。分布式机器学习(DML)方法(如联邦学习(FL))通过将数据和模型划分到多个服务器上来缓解这些挑战,但实现张量并行和流水线并行等优化仍然复杂。区块链技术作为一种有前景的解决方案应运而生,可确保分布式计算环境中的数据完整性、可扩展性和信任,但仍缺乏构建实用DML系统的指导。本文提出一种\textit{可信分布式机器学习}(TDML)框架,该框架利用区块链协调远程训练器并验证工作负载,在公共远程计算资源上实现隐私保护、透明且高效的模型训练。实验验证表明,TDML能有效克服性能限制并检测恶意节点,使其成为可扩展且安全的分布式机器学习的稳健解决方案。