This study proposes a trainable sampling-based solver for combinatorial optimization problems (COPs) using a deep-learning technique called deep unfolding. The proposed solver is based on the Ohzeki method that combines Markov-chain Monte-Carlo (MCMC) and gradient descent, and its step sizes are trained by minimizing a loss function. In the training process, we propose a sampling-based gradient estimation that substitutes auto-differentiation with a variance estimation, thereby circumventing the failure of back propagation due to the non-differentiability of MCMC. The numerical results for a few COPs demonstrated that the proposed solver significantly accelerated the convergence speed compared with the original Ohzeki method.
翻译:本研究提出一种基于深度展开这一深度学习技术的可训练采样求解器,用于解决组合优化问题。该求解器基于结合马尔可夫链蒙特卡洛与梯度下降的Ohzeki方法,并通过最小化损失函数来训练其步长参数。在训练过程中,我们提出一种基于采样的梯度估计方法,用方差估计替代自动微分,从而规避因马尔可夫链蒙特卡洛不可微性导致的反向传播失效。针对若干组合优化问题的数值实验表明,与原始Ohzeki方法相比,所提求解器显著加速了收敛速度。