With the availability of extraordinarily huge data sets, solving the problems of distributed statistical methodology and computing for such data sets has become increasingly crucial in the big data area. In this paper, we focus on the distributed sparse penalized linear log-contrast model in massive compositional data. In particular, two distributed optimization techniques under centralized and decentralized topologies are proposed for solving the two different constrained convex optimization problems. Both two proposed algorithms are based on the frameworks of Alternating Direction Method of Multipliers (ADMM) and Coordinate Descent Method of Multipliers(CDMM, Lin et al., 2014, Biometrika). It is worth emphasizing that, in the decentralized topology, we introduce a distributed coordinate-wise descent algorithm based on Group ADMM(GADMM, Elgabli et al., 2020, Journal of Machine Learning Research) for obtaining a communication-efficient regularized estimation. Correspondingly, the convergence theories of the proposed algorithms are rigorously established under some regularity conditions. Numerical experiments on both synthetic and real data are conducted to evaluate our proposed algorithms.
翻译:随着超大规模数据集的广泛应用,在大数据领域中解决此类数据的分布式统计方法与计算问题已变得日益重要。本文聚焦于大规模成分数据中的分布式稀疏惩罚线性对数对比模型。具体而言,针对两种不同的约束凸优化问题,分别提出了基于集中式与分散式拓扑结构的分布式优化技术。所提出的两种算法均以交替方向乘子法(ADMM)及乘子坐标下降法(CDMM,Lin等,2014,《生物计量学》)为框架。值得强调的是,在分散式拓扑结构中,我们引入了一种基于组ADMM(GADMM,Elgabli等,2020,《机器学习研究期刊》)的分布式坐标下降算法,以实现通信高效的规则化估计。相应地,在若干正则性条件下,严格建立了所提算法的收敛性理论。通过合成数据与真实数据的数值实验,对所提算法进行了评估。