Quantum computer simulators are crucial for the development of quantum computing. In this work, we investigate the suitability and performance impact of GPU and multi-GPU systems on a widely used simulation tool - the state vector simulator Qiskit Aer. In particular, we evaluate the performance of both Qiskit's default Nvidia Thrust backend and the recent Nvidia cuQuantum backend on Nvidia A100 GPUs. We provide a benchmark suite of representative quantum applications for characterization. For simulations with a large number of qubits, the two GPU backends can provide up to 14x speedup over the CPU backend, with Nvidia cuQuantum providing further 1.5-3x speedup over the default Thrust backend. Our evaluation on a single GPU identifies the most important functions in Nvidia Thrust and cuQuantum for different quantum applications and their compute and memory bottlenecks. We also evaluate the gate fusion and cache-blocking optimizations on different quantum applications. Finally, we evaluate large-number qubit quantum applications on multi-GPU and identify data movement between host and GPU as the limiting factor for the performance.
翻译:量子计算机模拟器对于量子计算的发展至关重要。在本工作中,我们研究了GPU及多GPU系统在广泛使用的模拟工具——状态向量模拟器Qiskit Aer上的适用性及性能影响。具体而言,我们在Nvidia A100 GPU上评估了Qiskit默认的Nvidia Thrust后端与最新的Nvidia cuQuantum后端的性能。我们提供了一套具有代表性的量子应用基准测试套件进行特性分析。对于大规模量子比特的模拟,两种GPU后端相比CPU后端可实现高达14倍的加速,而Nvidia cuQuantum相比默认的Thrust后端进一步提供1.5-3倍的加速。基于单个GPU的评估,我们识别了Nvidia Thrust和cuQuantum中针对不同量子应用的最关键函数及其计算与内存瓶颈。我们还评估了不同量子应用中的门融合与缓存阻塞优化。最后,我们在多GPU上评估了大规模量子比特的量子应用,并指出主机与GPU之间的数据传输是性能的制约因素。