The slowing down of Moore's law has driven the development of unconventional computing paradigms, such as specialized Ising machines tailored to solve combinatorial optimization problems. In this paper, we show a new application domain for probabilistic bit (p-bit) based Ising machines by training deep generative AI models with them. Using sparse, asynchronous, and massively parallel Ising machines we train deep Boltzmann networks in a hybrid probabilistic-classical computing setup. We use the full MNIST dataset without any downsampling or reduction in hardware-aware network topologies implemented in moderately sized Field Programmable Gate Arrays (FPGA). Our machine, which uses only 4,264 nodes (p-bits) and about 30,000 parameters, achieves the same classification accuracy (90%) as an optimized software-based restricted Boltzmann Machine (RBM) with approximately 3.25 million parameters. Additionally, the sparse deep Boltzmann network can generate new handwritten digits, a task the 3.25 million parameter RBM fails at despite achieving the same accuracy. Our hybrid computer takes a measured 50 to 64 billion probabilistic flips per second, which is at least an order of magnitude faster than superficially similar Graphics and Tensor Processing Unit (GPU/TPU) based implementations. The massively parallel architecture can comfortably perform the contrastive divergence algorithm (CD-n) with up to n = 10 million sweeps per update, beyond the capabilities of existing software implementations. These results demonstrate the potential of using Ising machines for traditionally hard-to-train deep generative Boltzmann networks, with further possible improvement in nanodevice-based realizations.
翻译:摩尔定律的放缓推动了非常规计算范式的发展,例如专为解决组合优化问题而设计的伊辛机。本文展示了基于概率比特(p-bit)的伊辛机的一个新应用领域——利用它们训练深度生成式AI模型。通过使用稀疏、异步且大规模并行的伊辛机,我们在混合概率-经典计算框架中训练深度玻尔兹曼网络。我们使用完整的MNIST数据集,未进行任何下采样或缩减,并在中等规模现场可编程门阵列(FPGA)上实现了硬件感知的网络拓扑结构。我们的机器仅使用4,264个节点(p-bit)和约30,000个参数,即可达到与优化后基于软件的受限玻尔兹曼机(RBM)相同的分类准确率(90%),而后者需约325万个参数。此外,稀疏深度玻尔兹曼网络还能生成新的手写数字,这是具有325万个参数的RBM尽管准确率相同却无法完成的任务。我们的混合计算机实测每秒可执行500至640亿次概率翻转,比表面类似的基于图形处理单元和张量处理单元(GPU/TPU)的实现至少快一个数量级。该大规模并行架构能轻松执行对比散度算法(CD-n),每次更新可进行多达1,000万次扫描,远超现有软件实现的能力。这些结果展示了利用伊辛机处理传统上难以训练的深度生成式玻尔兹曼网络的潜力,并在基于纳米器件的实现中具有进一步改进的可能。