节点文献

后量子密码算法ML-KEM快速数论变换的FPGA优化实现

FPGA-Based Optimized Implementation of Fast Number-Theoretic Transform of the Post-Quantum Cryptography Algorithm ML-KEM

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 高献伟; 李奕彤; 戴艺佳; 李天尧;

【Author】 GAO Xianwei;LI Yitong;DAI Yijia;LI Tianyao;Beijing Electronic Science and Technology Institute;

【机构】 北京电子科技学院;

【摘要】 在后量子密码标准算法ML-KEM中,基于数论变换的多项式乘法是核心计算瓶颈。实现低延迟、高吞吐量的NTT硬件加速器,必须同时解决计算单元的关键路径延迟与存储系统的带宽瓶颈两大挑战。为应对此挑战,本文提出了一种计算与存储协同设计的高性能FPGA实现方案。该方案的核心是一个由32路并行蝶形计算单元和32体分布式RAM阵列构成的、高度耦合的系统架构。其中,为时间抽取和频率抽取算法定制的5级深度流水线蝶形单元,将关键路径延迟降低,使系统在171 MHz的时钟频率下运行;而32体并行存储架构则提供了匹配高速计算核心所需的数据带宽,从架构上降低了数据的传输时间。该协同架构通过采用DIT/DIF混合算法优化了中间位反转开销,并利用资源高效的流水线化蒙哥马利模乘器,进一步提升了系统效率。在Zynq-7000 FPGA平台的实现结果表明,本设计需42个时钟周期0.245μs完成NTT变换,48个时钟周期0.281μs完成INTT变换,在执行速度和资源效率方面均优于部分方案,为后量子密码应用提供了一个可行的硬件加速方案。

【Abstract】 In the post-quantum cryptography standard algorithm ML-KEM, polynomial multiplication based on the Number-Theoretic Transform(NTT) is the core computing bottleneck. To achieve a high-throughput NTT hardware accelerator, it is necessary to simultaneously address two major challenges: the critical path latency of the computing unit and the bandwidth bottleneck of the storage system. To address these challenges, this paper proposes a high-performance FPGA implementation scheme featuring collaborative design of computing and storage. The core of the scheme is a highly coupled system architecture composed of 32 parallel butterfly computing units and a 32-bank distributed RAM array. Specifically, the 5-stage deeply pipelined butterfly unit, customized for decimation-in-time(DIT) and decimation-in-frequency(DIF) algorithms, reduces the critical path delay, enabling the system to operate at a clock frequency of 171 MHz. The 32-bank parallel memory architecture provides sufficient data bandwidth to match the high-speed computing core, thereby reducing data transfer latency at the architectural level. This collaborative architecture optimizes the intermediate bit-reversal overhead by adopting a DIT/DIF hybrid algorithm and further enhances system efficiency through a resource-efficient pipelined Montgomery modular multiplier. Implementation results on the Zynq-7000 FPGA platform show that the proposed design completes the NTT in 42 clock cycles(0.245 μs) and the inverse NTT(INTT) in 48 clock cycles(0.281 μs), which outperforms several existing schemes in terms of both execution speed and resource efficiency, providing a feasible hardware acceleration solution for post-quantum cryptography applications.

  • 【文献出处】 北京电子科技学院学报 ,Journal of Beijing Electronic Science and Technology Institute , 编辑部邮箱 ,2026年01期
  • 【分类号】O413;TN918.4
  • 【下载频次】34
节点文献中: