The hard mathematical problems that assure the security of our current public-key cryptography (RSA, ECC) are broken if and when a quantum computer appears rendering them ineffective for use in the quantum era. Lattice based cryptography is a novel approach to public key cryptography, of which the mathematical investigation (so far) resists attacks from quantum computers. By choosing a module learning with errors (MLWE) algorithm as the next standard, National Institute of Standard & Technology (NIST) follows this approach. The multiplication of polynomials is the central bottleneck in the computation of lattice based cryptography. Because public key cryptography is mostly used to establish common secret keys, focus is on compact area, power and energy budget and to a lesser extent on throughput or latency. While most other work focuses on optimizing number theoretic transform (NTT) based multiplications, in this paper we highly optimize a Toom-Cook based multiplier. We demonstrate that a memory-efficient striding Toom-Cook with lazy interpolation, results in a highly compact, low power implementation, which on top enables a very regular memory access scheme. To demonstrate the efficiency, we integrate this multiplier into a Saber post-quantum accelerator, one of the four NIST finalists. Algorithmic innovation to reduce active memory, timely clock gating and shift-add multiplier has helped to achieve 38% less power than state-of-the art PQC core, 4x less memory, 36.8% reduction in multiplier energy and 118x reduction in active power with respect to state-of-the-art Saber accelerator (not silicon verified). This accelerator consumes 0.158mm2 active area which is lowest reported till date despite process disadvantages of the state-of-the-art designs.
翻译:支撑当前公钥密码体系(RSA、ECC)安全性的数学难题一旦量子计算机问世即被破解,使其在量子时代失效。格密码作为一种新型公钥密码方法,其数学研究(迄今)能抵抗量子计算机攻击。美国国家标准与技术研究院(NIST)通过选择基于模误差学习(MLWE)算法作为下一代标准,正是遵循该路径。多项式乘法是格密码计算的核心瓶颈。由于公钥密码主要用于建立共享密钥,因此重点在于紧凑的面积、功耗与能量预算,吞吐量或延迟次之。当大多数研究聚焦于优化基于数论变换(NTT)的乘法时,本文高度优化了基于Toom-Cook的乘法器。我们证明采用延迟插值的存储高效步进Toom-Cook方法可实现高度紧凑、低功耗的实现方案,且能获得极规整的内存访问模式。为验证效率,我们将该乘法器集成至Saber后量子加速器(NIST最终入选的四项算法之一)。通过减少活跃存储器的算法创新、及时时钟门控及移位相加乘法器,相较最先进的PQC核心功耗降低38%,存储器减少4倍,乘法器能量降低36.8%,有功功率较最先进Saber加速器(未流片验证)降低118倍。本加速器占用0.158mm²有源面积,该面积值即便在工艺劣势下仍属迄今最低纪录。