The stochastic gradient descent (SGD) algorithm has been widely used to optimize deep Cox neural network (Cox-NN) by updating model parameters using mini-batches of data. We show that SGD aims to optimize the average of mini-batch partial-likelihood, which is different from the standard partial-likelihood. This distinction requires developing new statistical properties for the global optimizer, namely, the mini-batch maximum partial-likelihood estimator (mb-MPLE). We establish that mb-MPLE for Cox-NN is consistent and achieves the optimal minimax convergence rate up to a polylogarithmic factor. For Cox regression with linear covariate effects, we further show that mb-MPLE is $\sqrt{n}$-consistent and asymptotically normal with asymptotic variance approaching the information lower bound as batch size increases, which is confirmed by simulation studies. Additionally, we offer practical guidance on using SGD, supported by theoretical analysis and numerical evidence. For Cox-NN, we demonstrate that the ratio of the learning rate to the batch size is critical in SGD dynamics, offering insight into hyperparameter tuning. For Cox regression, we characterize the iterative convergence of SGD, ensuring that the global optimizer, mb-MPLE, can be approximated with sufficiently many iterations. Finally, we demonstrate the effectiveness of mb-MPLE in a large-scale real-world application where the standard MPLE is intractable.
翻译:随机梯度下降(SGD)算法已广泛用于优化深度Cox神经网络(Cox-NN),通过使用小批量数据更新模型参数。我们证明SGD旨在优化小批量部分似然的平均值,这与标准部分似然不同。这一区别需要为全局优化器(即小批量最大部分似然估计量,mb-MPLE)建立新的统计性质。我们建立了Cox-NN的mb-MPLE的一致性,并证明其达到了极大多项式因子内的最优极小化收敛速度。对于具有线性协变量效应的Cox回归,我们进一步表明mb-MPLE具有$\sqrt{n}$一致性且渐近正态,其渐近方差随着批量大小增加而趋于信息下界,仿真研究证实了这一点。此外,我们基于理论分析和数值证据提供了使用SGD的实践指南。对于Cox-NN,我们证明了学习率与批量大小的比值在SGD动态中至关重要,这为超参数调优提供了洞见。对于Cox回归,我们刻画了SGD的迭代收敛特性,确保全局优化器mb-MPLE可以通过足够多次迭代近似得到。最后,我们在标准MPLE难以处理的真实大规模应用中展示了mb-MPLE的有效性。