Generative machine learning methods such as large-language models are revolutionizing the creation of text and images. While these models are powerful they also harness a large amount of computational resources. The transformer is a key component in large language models that aims to generate a suitable completion of a given partial sequence. In this work, we investigate transformer architectures under the lens of fault-tolerant quantum computing. The input model is one where trained weight matrices are given as block encodings and we construct the query, key, and value matrices for the transformer. We show how to prepare a block encoding of the self-attention matrix, with a new subroutine for the row-wise application of the softmax function. In addition, we combine quantum subroutines to construct important building blocks in the transformer, the residual connection and layer normalization, and the feed-forward neural network. Our subroutines prepare an amplitude encoding of the transformer output, which can be measured to obtain a prediction. Based on common open-source large-language models, we provide insights into the behavior of important parameters determining the run time of the quantum algorithm. We discuss the potential and challenges for obtaining a quantum advantage.
翻译:大型语言模型等生成式机器学习方法正在彻底改变文本和图像的创建方式。尽管这些模型功能强大,但也消耗大量计算资源。Transformer是大型语言模型中的关键组件,旨在生成给定部分序列的合适补全。在本工作中,我们从容错量子计算的角度研究Transformer架构。输入模型以块编码形式给出训练好的权重矩阵,我们在此基础上构建Transformer的查询矩阵、键矩阵和值矩阵。我们展示了如何制备自注意力矩阵的块编码,并提出了逐行应用softmax函数的新子程序。此外,我们结合量子子程序构建了Transformer中的重要构建模块:残差连接与层归一化,以及前馈神经网络。我们的子程序可制备Transformer输出的振幅编码,通过测量即可获得预测结果。基于常见的开源大型语言模型,我们深入分析了决定量子算法运行时间的关键参数特性。最后,我们探讨了实现量子优势的潜力与挑战。