2463 results sorted by ID
DP-Guided Schedule Selection for $\mathbb{F}_2[x]$ Multiplication across ISAs
Junyu Zhou, Xiao Lan, Jing Wang, Hao Ren, Weiran Liu, Si Gao
Implementation
Efficient multiplication over $\mathbb{F}_2[x]$ is a core primitive in classical and post-quantum cryptographic software. As ARM and RISC-V become increasingly relevant for open-source cryptographic libraries such as OpenSSL and liboqs, arithmetic kernels must be retuned across a wider range of Instruction Set Architectures (ISAs). High-performance arithmetic libraries recursively apply Karatsuba- and Toom-style decomposition rules, each splitting the operands into smaller subproblems and...
Multiplying Not-So-Big Integers with FFTs: Finding and Pushing the Crossover on Large Modern OoO ARM CPUs
Cesare Huang, Daisy Meng-Wei Liu, David Shu-Yu Wu, Bow-Yaw Wang, Bo-Yin Yang
Implementation
Established multiple-precision software generally reserves Fast-Fourier Transform (FFT) multiplication for very large operands; for example, GMP reports full-product FFT thresholds of roughly 3000--10000 limbs. Becker et al. showed that Number-Theoretic Transform (NTT) multiplication can become advantageous much earlier on Cortex-M microcontrollers. It thus becomes an interesting engineering question to check where the crossover actually takes place on a big modern CPU. Our first-generation...
Comparative Evaluation of Open-Source Gröbner Basis Implementations for Small-Scale AES Polynomial Systems over GF(2)
Pavel Holý, Martin Jureček
Implementation
The practical performance of Gröbner basis implementations depends on the
polynomial systems being solved. This work evaluates 20 open-source imple-
mentations with Magma as a proprietary reference on systems derived from
small-scale AES over GF(2). The Easy, Medium, and Hard experiments use
field equations and ten input instances each, comparing wall-clock runtime, peak
memory, success rate, and CPU usage. Implementations with no successful runs
are excluded from subsequent...
Area-Efficient Hardware Architecture of the Post-Quantum HQC Cryptosystem with Side-Channel Protection
Kamal Raj, Peizhou Gan, Sayan Das, Anupam Chattopadhyay
Implementation
The Hamming-Quasi Cyclic (HQC) is the Post-Quantum Cryptography (PQC) Key Encapsulation Mechanism (KEM) algorithm recently selected by the National Institute of Standards and Technology (NIST) for standardization as an alternative candidate for the KEM category. While the algorithm is mathematically robust, several studies state that the physical implementation of HQC remains vulnerable to power-based side-channel attacks (SCA). Countermeasures for such SCAs primarily focus on masking and...
BOLT-FHE: An Efficient Unified Framework for GPU-based TFHE Bootstrapping via On-Chip Local Tiling Strategies
Yanren Chen, Fangyu Zheng, Guang Fan, Jiankuo Dong, Wenxu Tang, Tian Zhou, Jingqiang Lin, Jiwu Jing
Implementation
Bootstrapping is the main performance bottleneck in bitwise Fully Homomorphic Encryption (FHE), and practical acceleration requires careful orchestration of the blind rotation and external product chain under GPU resource constraints. This paper presents BOLT-FHE, a GPU bootstrapping framework that emphasizes block-local execution, on-chip tiling, and a unified MegaKernel supporting both gadget decomposition and modulus raising, with optional support for a recently proposed technique...
GPU-Assisted FAEST v3 Signing: Challenge Grinding and CPU–GPU Pipelining
Ha-Gyeong Kim, Si-Woo Eum, Seung-Won Lee, Min-Ho Song, Hwa-Jeong Seo
Implementation
Applying a GPU to FAEST v3 signing requires more than exploiting parallel computation: challenge grinding must preserve the minimum accepting counter selected by the original signing procedure, and single-sign latency and multi-sign throughput must be evaluated under distinct execution conditions. We implement GPU-assisted challenge grinding in which the GPU filters candidates over contiguous counter ranges in parallel, while the CPU performs the final acceptance checks in increasing counter...
X-Wing on Cortex-M85: Architecture-Aware Integration and End-to-End Evaluation
Do-Yun Park, Hwa-Jeong Seo
Implementation
X-Wing combines ML-KEM-768 and X25519 for hybrid key establishment. On Cortex-M85, we adapted the coefficient representation and storage order of an MVE NTT to the existing ML-KEM path and optimized data-format transformations and X25519 arithmetic while preserving draft-10. During encapsulation, the two X25519 results share an inversion. We compared the eight-technique integrated implementation B8 with the baseline A0 in the same executable. Across ten paired comparisons from five...
Batched LSH on ARMv8 NEON and Its Application to K-SPHINCS+
Ji-Won Bang, Su-Been Cho, Yu-Lim Hyoung, Ha-Gyeong Kim, Hwa-Jeong Seo
Implementation
With FIPS 205 standardizing post-quantum signatures and ARMv8 spanning mobile, embedded, and server platforms, fast hash-based signing on ARM has become a pressing need. Since these schemes are dominated by hash calls, hash optimization largely determines performance. K-SPHINCS+ replaces the internal hash of stateless SPHINCS+ with LSH, a Korean standard hash family. However, prior K-SPHINCS+ work provides only reference C code, and existing ARMv8 LSH implementations mainly optimize...
Optimized NEON Implementation of MAYO on Apple M1
Minwoo Lee, Minjoo Sim, Siwoo Eum, Hwajeong Seo
Implementation
MAYO is one of nine third-round candidates in NIST’s process for additional post-quantum signatures. Its third-round specification of August 2026 changed the level-1 parameters and added a hash-derived linear term Λ to the verification equation, so earlier optimised implementations no longer match it. We present an optimised NEON implementation of MAYO Round 3 for the Apple M1 and compare it with the upstream NEON code on the same machine. A generator emits the sixteen GF(16) kernels of each...
Verifier-Aligned Canonical Source Binding for Homomorphic Encryption Computations
Subeen Cho, Jiwon Bang, Minseo Kim, Seungwon Lee, Hwajeong Seo
Implementation
q0 preprocessing verification and native HE computation verification can each succeed without establishing that their results refer to the same canonical source. To address this source mismatch across verification paths, we propose a verifier-aligned public-input transfer procedure that links heterogeneous authenticated artifacts to a single canonical-source relation. The public verifier reconstructs challenges and targets from the actual statements, roots, approved verification keys, and...
HyBind: Structural Analysis of PQ/T Hybrid Cryptographic Usage
Jongbeom Ahn, Minseo Kim, Moonsun Heo, Subeen Cho, Hwajeong Seo
Implementation
The quantum threat to traditional public-key cryptography makes migration to post-quantum cryptography necessary. During migration, the presence of both post-quantum and traditional (PQ/T) algorithms does not establish that their outputs jointly determine a key or verification decision. We propose HyBind, an intraprocedural Python source-analysis framework using registered API contracts. HyBind traces component secrets to returned keys and checks whether signature acceptance requires both...
Randomness-Anchored Runtime Discovery of Cryptographic Operations on Linux Using eBPF
Sumin Jeong, Yulim Hyoung, Hagyeong Kim, Huiju Kang, Hwajeong Seo
Implementation
Migration to post-quantum cryptography requires knowing not only which cryptographic libraries are present on a system, but which algorithms are actually executed at runtime. Static binary analysis answers the former question; the latter requires dynamic observation. We present a runtime cryptographic discovery system for Linux, built on eBPF user-space probes, that detects classical and post-quantum operations as they execute. Because key generation and encapsulation consume fresh...
Role-Specific Signature Placement in TLS 1.3: Cost Propagation and PQC Migration
Moonsun Heo, Subeen Cho, Jongbeom Ahn, Minseo Kim, Hwajeong Seo
Implementation
Certificate issuance and per-connection signing impose different costs in TLS 1.3, and authentication cost can change depending on how signature algorithms are assigned to Root, Intermediate, and Leaf roles. PQC migration should therefore consider not only algorithm performance, but also role placement, certificate structure, network conditions, and server load. We evaluate all 13³ role placements formed by 13 signature parameter sets while holding the algorithm composition fixed and...
Post-Quantum Command Authentication for Unmanned Systems: A KpqC Feasibility Study under Real-Time and Bandwidth Constraints
Yulim Hyoung, Daeun Lim, Jiwon Bang, Doyun Park, Hagyeong Kim, Hwajeong Seo
Implementation
The command and authentication channels of unmanned systems (UAVs and UGVs) still rely on classical public-key cryptography, which large-scale quantum computers threaten. Migrating these channels to post-quantum cryptography (PQC) is therefore imperative, yet PQC keys and signatures are substantially larger, and sometimes slower, than their classical counterparts, while command links are tightly constrained in message rate, maximum transmission unit (MTU), latency deadline, and bandwidth....
Deploying Native-Rust KpqC: Pluggable Post-Quantum KEMs and Signatures across TLS, Space Links, and the OQS Ecosystem
Yulim Hyoung, Doyun Park, Hyunji Kim, Hwajeong Seo
Implementation
Post-quantum migration is no longer only an algorithm problem but a deployment problem: standardized schemes must be dropped into real protocol stacks, on real (often resource-constrained) targets, without giving up memory safety. The Korean Post-Quantum Cryptography (KpqC) competition standardized two KEMs (NTRU+, SMAUG-T) and two signature schemes (HAETAE, AIMer), distributed as memory-unsafe C. Building on our native-Rust KpqC implementation, we expose the schemes through a single...
Measuring Parallel Execution of Classical–PQC Component Pairs in Hybrid TLS 1.3
Siwoo Eum, Minho Song, Seung-Won Lee, Hwajeong Seo
Implementation
Hybrid key exchange and Composite ML-DSA signatures in TLS 1.3 split one operation into a classical component and a post-quantum cryptography (PQC) component that do not depend on each other. Using a cryptographic workload of the TLS handshake rather than a full TLS connection, this paper measures whether running the two components at the same time on two pthread worker threads reduces latency compared with sequential execution. We evaluate the three ECDHE–ML-KEM groups of RFC 10024 with the...
Optimizing Montgomery Arithmetic for RSA on Cortex-M0+ and Cortex-M3
Minoo Sim, Minwoo Lee, Seungwon Lee, SuBeen Cho, Jiwon Bang, Hwajeong Seo
Implementation
Cryptographic tokens that retain RSA credentials for compatibility with existing authentication systems require efficient private-key operations on small processors. Cortex-M0+ (M0+) lacks native widening multiplication, whereas Cortex-M3 (M3) provides it with operand-dependent latency. We optimize Montgomery multiplication and squaring using 15-bit limbs and integrate the kernels into RSA-2048 and RSA-3072 private operations based on the Chinese remainder theorem (CRT). On M0+, a wrapped...
What Makes Lattice Key Generation Expensive? Controlled Cost Attribution with Structured LWR, ML-KEM, and HAETAE on Cortex-M4
Yan Zhang, Meizi Li, Liang Tan
Implementation
End-to-end cycle counts quantify lattice key-generation time on a microcontroller, but not which implementation decisions create that cost. We develop a controlled attribution method for public structure, candidate admission, and transform lifetime: how public data are organised, when candidate acceptance is checked, and whether transformed secret state is retained or reconstructed. On a fixed STM32L476RG Cortex-M4 target, paired runs within one executable keep the relevant secret, accepted...
ECLIPSE: Strongly Unforgeable Isogeny Signatures from the Prime-Degree Variant of PRISM
Dustin Ray
Implementation
ECLIPSE is the prime-degree signature construction that the PRISM authors describe next to their salted scheme, with that salt, implemented. It exists because a two-dimensional isogeny signature carries an auxiliary isogeny that the hash does not bind: the SQIsign round-3 specification states that SQIsign cannot achieve strong unforgeability for this reason, and PRISM's prime-degree variant was designed without the auxiliary isogeny but published without parameters, code or a verifier that...
Low-Noise Multi-Value Bootstrapping via Log-Unit Lattice Search and LUT Shifting
Olivier Bernard, Nolan Carouge, Marc Joye, Jean-Baptiste Orfila, Samuel Tap
Implementation
Programmable bootstrapping is a central procedure in the FHEW and TFHE families of fully homomorphic encryption schemes, but its computational cost remains a major performance bottleneck. Multi-value bootstrapping amortizes this cost by sharing a blind rotation among several functions evaluated on the same encrypted input. However, the standard choice of common factor for multi-value bootstrapping can produce unnecessarily large noise amplification for many collections of lookup tables,...
Functional Verification of Additive FFT Assembly Implementations in Code-Based PQC
Jiaxiang Liu, Kuang-Lin Pan, Xiaomu Shi, Ming-Hsien Tsai, Bow-Yaw Wang, Bo-Yin Yang
Implementation
We present the first formal verification results for the highly optimized additive Fast Fourier Transforms (FFTs) in two different code-based post-quantum cryptosystems: Classic McEliece and Hamming Quasi-Cyclic. Since Classic McEliece is already a standard and HQC is being standardized, they may both see wide adoption in the future. The Gao-Mateer and Frobenius additive FFTs during decoding are clearly the most intricate, error-prone components in the two cryptosystems respectively. They...
Knuth–Yao masked (Gaussian) sampling
Calvin Abou Haidar, Clément Hoffmann
Implementation
Discrete Gaussian sampling remains one of the most delicate operations to protect against side-channel attacks in lattice-based cryptography. In this work, we present the first masked evaluation of the Knuth-Yao sampler, a generic building block for lattice-based schemes. Its random-walk formulation might seem fundamentally at odds with masking, since the walk's control flow depends precisely on the secret sample being produced. We show instead that its underlying tree structure is...
Masked CROSS: Masking the CROSS Digital Signature Scheme at Arbitrary Order
Khan Keren Mengwi, Puja Mondal, Achille Ecladore Tchahou Tchendjeu, Emmanuel Fouotsa, Suparna Kundu
Implementation
CROSS is a code-based signature scheme built on the Restricted Syndrome Decoding Problem and is a second-round candidate in NIST's additional digital signature standardization process. Recent work shows that its reference implementation is vulnerable to side-channel analysis, allowing an attacker to recover the long-term secret key from a single power trace. No masking countermeasure has so far been designed for CROSS's restricted-syndrome framework, leaving its practical side-channel...
ML-DSA masking sweetened with SUCRE: Shuffle-and-Unmask Countermeasure for REjection sampling
Sonia Belaïd, Ryad Benadjila, Julien Devevey, Morgane Guerreau, Thomas Legavre, Ange Martinelli, Thomas Ricosset, Matthieu Rivain, Mélissa Rossi
Implementation
We present SUCRE, a novel countermeasure designed to physically protect the rejection sampling step of ML-DSA, one of the post-quantum signature schemes standardized by NIST. At the core of SUCRE is a masking gadget that securely unmasks a vector while simultaneously applying a random permutation of its coefficients. This lightweight mechanism preserves the vector’s infinity norm, enabling rejection sampling to proceed as usual without requiring any complex mask conversions.
We formally...
Dimension-4 SQIsign at Round-3 Parameters: Sizes and Costs of the Compact Format
Dustin Ray
Implementation
SQIsign holds the smallest known combination among post-quantum signatures, and its dimension-2 signature carries an auxiliary curve whose only role is to make the response computable in dimension 2. SQIsignHD's dimension-4 response has no such curve. We take that construction to the round-3 SQIsign primes, which replaced the round-2 primes in September 2026 after the attacks of Wesolowski and others, on our Rust port of the SQIsign reference code, and measure it in one session next to the...
A New Prime-Norm Ideal Sampling for SQIsign
Li-Jie Jian
Implementation
SQIsign samples a random ideal of odd prime norm at every key generation and every signature through $\mathbf{PrimeNormIdealInertSampling}$. We revisit this step from a new perspective: the inert ideals of a given norm form a circle modulo the norm, and drawing lines through a fixed point of this circle turns a single random parameter into an ideal, hitting each ideal exactly once.
Compared with the reference implementation, the new algorithm requires one modular inversion and a handful...
Hardware-Adapted SIMD Optimization of AIMer v3 on x86 and Cortex-M55
Seung-Won Lee, Si-Woo Eum, Min-Seo Kim, Su-Been Cho, Hwa-Jeong Seo
Implementation
AIMer v3 replaces AIM2 with AIM3. We test whether SIMD methods developed for AIM2 can accelerate AIM3 while preserving its results. We port the field-arithmetic, affine, party-parallel, and SHAKE methods of the earlier AVX-512 implementation to AIMer v3 and implement the same operations for AVX2. For Cortex-M55 MVE, we use 16-bit polynomial partial products for field multiplication, group corresponding 32-bit words from four parties, and vectorize both affine layers, including the input...
High-Performance Implementations of HAETAE and SMAUG-T on Arm Cortex-M55 Using MVE and Slothy
Seung-Won Lee, Min-Joo Sim, Hwa-Jeong Seo
Implementation
Efficient execution of Post-Quantum Cryptography (PQC) in resource-constrained embedded environments requires the joint optimization of not only algorithm vectorization but also dataflow and instruction execution order. This paper implements the HAETAE-2/3/5 and SMAUG-T1/3/5 parameter sets of two Korean Post-Quantum Cryptography Competition (KpqC) winners using the M-profile Vector Extension (MVE) of the Arm Cortex-M55 and analyzes the contribution of each optimization stage to performance...
New algorithms for quaternion ideals in SQIsign
Antonin Leroux, Sina Schaeffler
Implementation
Many isogeny-based schemes rely on the Deuring correspondance and thus require computations with quaternion ideals.
In this paper, we generalize the recent results of Leroux on the inert representation of quaternion ideals, yielding a normalized representation together with a set of simple and efficient algorithms for quaternion ideals in the context of the Deuring correspondence when the prime characteristic $p$ is equal to $3 \bmod 4$.
One of the main benefit of our new algorithms...
Too Small to Hide: Single-Trace Key Recovery from ML-KEM Key Generation
Vahid Jahandideh
Implementation
Key generation in ML-KEM (CRYSTALS-Kyber) samples a short secret from a centered binomial distribution (CBD) and immediately transforms it with the number-theoretic transform (NTT). Each execution draws fresh randomness, so an attacker obtains a single trace and cannot average. We show that one power trace of the optimized pqm4 implementation on an Arm Cortex-M4 suffices, and that the two operations leak complementary information. The CBD sampler stores each coefficient as a signed $16$-bit...
Fast Quantum-Circuit Superoptimization
Aws Albarghouthi
Implementation
Optimizing quantum circuits is critical: circuits must fit within the resource limits of a quantum computer, and every unnecessary operation increases their cost and probability of failure.We present a simple optimization algorithm for quantum circuits that (1) is very fast, (2) scales to millions of operations, and (3) matches or outperforms the optimization quality of the best existing optimizers and superoptimizers.
Our key insight is that we can compactly represent circuit equivalence...
Computing C(N, R) mod m for Arbitrary Composite Moduli and N <= 10^18: A Prime-Power Decomposition Engine with Two Memory Regimes and Division-Free Arithmetic
Vadik Malik, Rudr Pratap, Sarthak Vashishtha
Implementation
Binomial coefficients modulo an integer, $\binom{N}{R} \pmod m$, are a primitive of combinatorial counting, yet the two textbook methods collapse at scale: the Pascal recurrence costs $\Theta(NR)$ time, and the factorial-table method costs $\Theta(N)$ memory, requires a prime modulus, and requires $N < m$. Composite moduli are harder still, because factorials are not invertible modulo prime powers and the exact power of $p$ dividing the coefficient must be tracked. We present a complete,...
Quantum implementation of Keccak-f(25) for emulator and quantum hardware, and application to quantum password cracking
Mathilde Chenu, Mario Chizzini
Implementation
In this work, we present a quantum implementation of Keccak-$f(25)$, a toy version of SHA-3 introduced in the specification of Keccak, with the objective of using as few qubits as possible so that the resulting implementation can be run both on quantum emulators and on quantum hardware.
The code, written in Qiskit, uses $25$ qubits representing the internal state. All computations are performed in-place, meaning that no ancillary qubit is required.
We use this implementation to...
Optimizing HAETAE and SMAUG-T on Cortex-M4 for Resource-Constrained IoT Devices
JunHyeok Choi, DongHyun Shin, Seog Chung Seo
Implementation
The practicality of post-quantum cryptography (PQC) on IoT devices depends on both operation cycle counts and communication costs from public values (public keys, ciphertexts, and signatures). The DSA HAETAE and the KEM SMAUG-T, both selected in the KpqC competition, provide smaller public values than ML-DSA and ML-KEM; however, the lack of platform-specific optimization leaves their operation cycle counts high, which can offset this advantage. In this paper, we optimize HAETAE and SMAUG-T...
Area-Time Efficient NTRU Prime Decapsulation: ASIC Evaluation of the First Five-Way Char-3 Multiplier
Esra Yeniaras
Implementation
Streamlined NTRU Prime (sntrup761) is a lattice-based key encapsulation mechanism that, although not a NIST standard, remains widely deployed in critical internet infrastructure. It is the post-quantum key-exchange default in OpenSSH, standardized in RFC 9941, and used well beyond SSH, in Red Hat Enterprise Linux, the liboqs library, PQConnect, and commercial VPNs. Its decapsulation performs a polynomial multiplication over the characteristic-three ring $\mathbb{Z}_3[x]/(x^{p}-x-1)$, so...
Trace-Factored BigSwitch for Matrix-Friendly FHE
Dong Jin Park, Hyunseok Jeong, Minwook Jeong, Jaeky Oh, Yongwoo Lee, Young-Sik Kim
Implementation
Evaluation keys represent a primary memory and initialization bottleneck in matrix-native fully homomorphic encryption (FHE). In the Gentry–Lee (GL) framework, each Trace product yields a four component ciphertext whose BigSwitch procedure requires two extended-ring keys, dominated by a massive product-secret key (sXsY → sX). We present Trace-Factored BigSwitch (TFB), which structurally eliminates this product-secret evaluation key by exploiting the rank one tensor structure of the...
Square Root of All Evil: The Dangers of Falcon's Superfluous Square Roots
Hiroto Kaihara, Calvin Abou Haidar, Mehdi Tibouchi, Masayuki Abe
Implementation
Falcon is one of the 3 post-quantum signature schemes already selected by NIST for standardization (as FN-DSA). It is very compact and efficient, but also infamously difficult to implement correctly and securely. This is due in particular to its reliance of various floating point operations, the most complex and costly of which are square root computations.
In this paper, we first point out that those square root computations are in fact wholly unnecessary: the algorithm can be rewritten...
A Unified Analysis of Refresh Gadgets in the Random Probing Model
Sonia Belaïd, Victor Normand, Matthieu Rivain
Implementation
Masking is a standard countermeasure against side-channel attacks on embedded cryptographic implementations. Its security is commonly analyzed in the random probing model, which offers a useful trade-off between realistic leakage assumptions and tractable security proofs. Recent years have seen the emergence of several masking compilers based on compositional security frameworks such as general/cardinal random probing composability (RPC). Most of these approaches rely on dedicated refresh...
Beasley: Efficient Zero-Knowledge Proofs for Lattice-based Round-Optimal Oblivious Pseudorandom Functions
Baiyu Li, Ting-Yuan Wang, Jiapeng Zhang
Implementation
Verifiable oblivious pseudorandom functions (VOPRFs) enable a client to evaluate a pseudorandom function on a private input under a server-held key while verifying that the server evaluated the function honestly using its committed key. Existing practical VOPRFs rely mainly on assumptions that are vulnerable to quantum attacks. Prior lattice-based constructions achieving both round optimality and malicious security have remained theoretical proposals without concrete implementations, largely...
Bounded Information: PAC Certification of Multivariate Side-Channel Traces
Kuheli Pratihar, Nimish Mishra, Debdeep Mukhopadhyay
Implementation
Side-channel leakage certification aims to quantify what an attacker can learn about a secret variable from observed leakage. Existing information-theoretic estimators, such as perceived information (PI), hypothetical information (HI), and nonparametric mutual information (MI), aim to quantify distributional leakage, but they become unstable in high-dimensional traces and do not provide a finite-sample certificate of the best attacker. We introduce \emph{Bounded Information} (BI), a probably...
A RAM-Efficient Implementation of Falcon
Thomas Pornin
Implementation
We present a RAM-efficient implementation of Falcon: RAM usage has shrunk to about 11 kB, down from about 31 kB in the previous implementation of Falcon-512. This code is furthermore faster on Arm Cortex M4, with average signature generation cost down to 13.45 million cycles. Optimization techniques include a novel variant of the FFT, replacement of some floating-point operations with modular integer computations, delayed addition of input within the Fast Fourier sampling process, and an...
Optimal Bucket Set Construction for Multi-scalar Multiplication with Endomorphisms
Nam Hoai Le, Francesco Sica
Implementation
We develop the method of Luo, Fu and Gong (LFG) - as extended by Fan, Kuchta, Sica and Xu (FKSX) to use endomorphism scalars - in order to find best families suitable for multi-scalar multiplication (MSM) for any number $n$ of points and with adjustable storage.
In particular we lower storage requirements by an average of 67% and decrease the number of curve operations by an average of 7% (and up to 10.6%), relative to the LFG method. Compared to Pippenger's variant (standard when $n$ is...
Constant-Time Conditions for Left-to-Right Scalar Multiplication
Sergey Agievich
Implementation
We consider methods for scalar multiplication on an elliptic curve where the scalar digits are processed from left to right, that is, from most significant to least significant. We analyze exceptions that may arise during point addition and doubling throughout the multiplication. Eliminating such exceptions is critical for achieving constant-time execution and preventing timing attacks. We establish conditions under which exceptions occur only at the final addition or not at all. Guided by...
PEEV: Parse Encrypt Execute Verify - A Verifiable FHE Framework
Omar Ahmed, Charles Gouert, Nektarios Georgios Tsoutsos
Implementation
Cloud computing has been a prominent technology that allows users to store their data and outsource intensive computations. However, users of cloud services are also concerned about protecting the confidentiality of their data against attacks that can leak sensitive information. Although traditional cryptography can be used to protect static data or data being transmitted over a network, it does not support processing of encrypted data. Homomorphic encryption can be used to allow processing...
Terrazzo: Memory-Aware GPU Framework for Private Transformer Inference
Rostin Shokri, Nektarios Georgios Tsoutsos
Implementation
Fully homomorphic encryption (FHE) allows a server to run inference directly on encrypted data, making it a promising foundation for private transformer inference. Its dominant scheme, CKKS, has no native matrix multiplication, so the encrypted matrix multiplications at the heart of transformers dominate inference cost. The recent GL scheme supports matrix multiplication natively, but its ciphertexts, plaintexts, and evaluation keys are so large that a direct GPU implementation would require...
Jacobian Diagnostics for Under-Constrained Zero-Knowledge Circuits
Vijay Singh
Implementation
Under-constrained arithmetic circuits are a recurring source of soundness failures in zero-knowledge applications: after fixing the public statement, a malicious prover may be able to assign a security-relevant wire in more than one way while still satisfying the circuit. Existing tools attack this uniqueness question with solver-based checking, direct polynomial solving, abstract interpretation, or fuzzing. We study a complementary algebraic diagnostic based on exact Jacobian linear...
WeaveTLS: High-Throughput Cross-Connection ML-DSA Authentication in Mutual TLS
Ganqin Liu, Hao Cheng, Jipeng Zhang
Implementation
Mutual TLS (mTLS) authenticates both peers and therefore incurs post-quantum
signature costs on every connection. Concurrent handshakes expose independent
ML-DSA operations, but executing them jointly is difficult: signing is
rejection-divergent, verification uses heterogeneous keys, and synchronous TLS
APIs expose authentication work one connection at a time.
We present WeaveTLS, a wire-transparent architecture that executes
ML-DSA authentication across concurrent TLS connections....
AVXPoS: Reducing Consensus Verification Cost in the Ethereum Proof-of-Stake Client
Ganqin Liu, Hao Cheng, Georgios Fotiadis, Jipeng Zhang, Chen Qian
Implementation
Ethereum Proof-of-Stake (PoS) clients must verify large volumes of
Boneh--Lynn--Shacham (BLS) signatures for attestations, sync-committee messages,
and other consensus-critical objects within fixed slot deadlines. This recurring
cost competes with state transition, fork choice, and message propagation for
client CPU time, so reducing it increases the verification headroom available
under bursty load. Prior cryptographic-engineering work has shown that SIMD can
substantially accelerate...
Novel SMT Encoding for Quantum Circuit Optimization
Youbo Guo, Fengrong Zhang, Lei Liao, Yongzhuang Wei, Baocang Wang, Xiaogang Zhou
Implementation
In recent years, quantum circuit optimization has become an important research topic. Motivated by the fact that quantum gates act on fixed physical wires and modify only their target wires, we propose two SMT encodings: an exact-G encoding and an at-most-G encoding with null gates. Our method speeds up most tested 4-bit S-box instances, achieving up to approximately 130x speedup on the ELEPHANT S-box. Importantly, our method enables automated synthesis of practical 5-bit S-box quantum...
Lithium: Making Iterative Rejection Sampling Practical for Compact Lattice Signatures
Jipeng Zhang, Pengfei Chen, Long Chen, Cong Zhang, Jiaheng Zhang
Implementation
Post-quantum deployments need signatures that are both fast and small. ML-DSA gives a practical Fiat-Shamir lattice-signature baseline, but its signatures remain large enough to make bandwidth, certificate size, and signed-log storage first-order costs. Gaertner's iterative rejection sampling construction (CRYPTO'25) shows that this design family can be made much more compact. The open question is whether this theoretical design can be turned into a concrete, implementation-oriented...
Faster Post-Quantum zkSNARK Provers Using the LCH Polynomial Basis
Mohammadtaghi Badakhshan, Susanta Samanta, Guang Gong
Implementation
Univariate-polynomial interactive oracle proofs (IOPs) over binary extension fields $\mathbb{F}_{2^m}$ underpin a class of plausibly post-quantum zkSNARKs, but rely heavily on polynomial arithmetic, where large-domain evaluation and division by subspace vanishing polynomials are the dominant prover costs. General-basis additive FFTs, such as Gao--Mateer and Lin-Chung-Han (LCH), accelerate the evaluation but impose a basis-conversion stage costing $O(n (\log n)^2)$ field additions and $O(n...
Two Novel Multidimensional Affine Variations of the Hill Cipher
Porter Eldridge Coggins
Implementation
Two novel symmetric multidimensional affine nested variations of the Hill Cipher are presented. The Hill Cipher is a block
polygraphic substitution encryption scheme based on a linear transformation of plaintext characters into ciphertext characters. In
the time since Hill first published his encryption scheme, variations, modifications, and improvements of theoretical and
practical importance have been published every year indicating that the Hill Cipher is an active area of cryptography...
Notes on Short-Limb Modular Multiplication Techniques: Barrett, Montgomery, Plantard, and the Explicit CRT
Bo-Yin Yang
Implementation
This note collects, in compressed form, some techniques for modular multiplication with
word-size (“short-limb”), or at most a-handful-of-words sized moduli as they are used in
implementations of lattice-based cryptography: Barrett reduction and multiplication (in
signed and unsigned flavors, with exact error, range, and canonicality analyses), Montgomery
reduction and multiplication (including the folded-constant form, the precise equivalence with
Barrett multiplication, even moduli,...
Adapting AES-Oriented Optimizations to Rijndael-256: Cortex-M4, ARMv8-A, and CUDA
Siwoo Eum, Minho Song, Minjoo Sim, Anupam Chattopadhyay, Hwajeong Seo
Implementation
Rijndael-256 (R256), the 256-bit block variant of the Rijndael family, is practically relevant in ongoing NIST draft discussions on wider-block standardization and in several NIST post-quantum signature candidates. Relative to AES, R256 combines a wider $4\times8$ state with non-standard ShiftRows offsets $(0,1,3,4)$, invalidating key assumptions behind many AES-oriented optimizations. We study how these mismatches manifest on three targets and develop three corresponding adaptation...
Exposing SIMD Parallelism in SQIsign: An AVX-512 Implementation
Weize Wang, Chutong Wang, Yu Wu, Qifan Xue, Jieyu Zheng, Yunlei Zhao
Implementation
Modern isogeny-based cryptosystems spend much of their running time in finite-field, elliptic-curve, and higher-dimensional isogeny arithmetic. Exploiting SIMD parallelism in these computations is nevertheless nontrivial: central routines such as Montgomery ladders contain loop-carried dependencies, while point, pairing, and theta-coordinate formulas expose only irregular fine-grained parallelism. We show that substantial SIMD parallelism can be recovered by reorganizing the arithmetic...
MAYO Lite: a Low-RAM Implementation of the MAYO Signature Scheme
Sven Bauer, Fabrizio De Santis, Florian Wilde
Implementation
MAYO is a signature scheme based on the Unbalanced Oil and Vinegar
(UOV) construction and a third round candidate in the NIST standardization process
for additional post-quantum signature schemes. We present a memory-optimized pure-
C implementation of MAYO signature verification that reduces RAM consumption
by 97–99% compared to the reference implementation provided by the PQM4 project
[KPR+] at the cost of increasing runtime by 50–200% and while maintaining code
size. This reduction...
Code Generation of Faster Formally Verified NTT with Plantard Reduction
Donnie Y. Xu, Rajeev Gore, Amin Sakzad, Ron Steinfeld, Raymond K. Zhao
Implementation
We present a formally verified implementation of the ML-KEM Number-Theoretic Transform (NTT) based on Plantard arithmetic, produced via a code generator that targets ML-KEM, ML-DSA, and FN-DSA from a single parameter triple. The generator embeds a static bound analyzer that places modular reductions at code-generation time without runtime branching, eliminating per-scheme manual tuning while preserving constant-time guarantees. Each generation produces structurally identical implementations...
LFSRs and Boolean Masking: An In-depth Security Analysis
Anna Guinet, Jan Schoone, Niklas Höher, Dina Hesse, Tim Güneysu
Implementation
Masking is a widely adopted countermeasure to protect cryptographic implementations from side-channel attacks. Subsequent research has focused on designing masking schemes and formally proving their security, notably through the development of automated tools, within models abstracting the reality of a sidechannel analysis. These designs rely on an external source of randomness; however, there is currently no consensus on the choice of (pseudo-)random number generators for masking. To the...
Algorithmic Optimization of the Gaussian Sampler in the FN-DSA Post-Quantum Signature Scheme
Nicolas HOULÈS, Thibaut Heckmann
Implementation
The post-quantum signature scheme Falcon (FN-DSA), currently being standardized by NIST as FIPS 206 (Initial Public Draft submitted August 2025, final standard expected 2026-2027), relies on a discrete Gaussian sampler whose critical bottleneck is the function fpr_expm_p63, computing $\lfloor \exp(-x) \cdot 2^{63} \rfloor$ for $x \in [0, \ln 2)$. While the reference implementation already employs a degree-12 fixed-point polynomial (FACCT), no segmented approximation has been studied for this...
Efficient Large-Integer Arithmetic for FHE
Ahmad Al Badawi, Andreea Alexandru, Gurgen Arakelov, Charles Gouert, Sergey Gomenyuk, Valentina Kononova, Yarkın Doröz, Yuriy Polyakov
Implementation
Fully Homomorphic Encryption (FHE) has emerged as one of the key technologies for privacy-preserving computation, enabling arbitrary computation directly on encrypted data. Vectorized FHE schemes, such as Brakerski/Fan--Vercauteren (BFV), Brakerski--Gentry--Vaikuntanathan (BGV), and Cheon--Kim--Kim--Song (CKKS), are typically used in applications dealing with large datasets, for example, confidential database queries and private ML inference. These FHE schemes are based on the computational...
Strided Frobenius Additive FFT and its Application to HQC
Ming-Shing Chen, Tun-You Chien, Chun-Ming Chiu, Cesare Huang, Han-Hsuan Lin, Chun-Tao Peng, Bo-Yin Yang
Implementation
Boolean polynomial multiplication is the primary computational bottleneck of the Hamming Quasi-Cyclic (HQC) key encapsulation mechanism. In this paper, we reframe the Frobenius Additive FFT (FAFFT) in ring-theoretic terms, via quotient-ring homomorphisms and the Chinese Remainder Theorem. This perspective shows that a complete decomposition into evaluation points is unnecessary for multiplication, and naturally yields the Strided FAFFT (SFAFFT), which operates over smaller finite fields with...
Beyond Affine Invariants: A Hamming-Weight Correlation Metric for Template-CPA Leakage in Key-Dependent S-boxes
Wiesław Maleszewski
Implementation
Classical selection criteria for cryptographic S-boxes—nonlinearity $\mathrm{NL}$, differential uniformity $\delta$, boomerang uniformity $\beta_{\mathrm{B}}$, algebraic degree $\deg$—are invariants of affine equivalence. That property is exactly what blinds them to a class of side-channel weaknesses. The correlation-power-analysis (CPA) template distinguisher is governed by the Hamming-weight functional, and Hamming weight is not affine-invariant; it does not descend to the...
A Systematic Literature Review on Optimising CRYSTALS-Dilithium (ML-DSA) Performance for IoT Devices via Lightweight Hashing
Ceasar Njuguna Ngunu, Edward Ombui
Implementation
Background: The migration to post-quantum cryptography confronts resource-constrained Internet of Things (IoT) devices with a material performance cost. CRYSTALS-Dilithium, standardised as the Module-Lattice-Based Digital Signature Algorithm (ML-DSA) in FIPS 204, fixes the Keccak-based SHAKE functions as its only symmetric primitives, and profiling on embedded platforms identifies hashing as the largest single contributor to the scheme’s software cost. This review synthesises the performance...
SHARMONY: Composing SHA-2 and SHA-3 Hardware for Crypto-Agile PQC
Liga Anwar, Carlos Andres Lara-Nino, Jong-Yeon Park, Michael Hutter
Implementation
This work composes SHA-2 and SHA-3 into a unified hardware architecture, bringing them together as a single, efficient cryptographic ensemble. This need is driven in particular by Post-Quantum Cryptography (PQC), where different standardized schemes rely on either SHA-2 or SHA-3/SHAKE primitives. Rather than enforcing strict round-level unification, the proposed design applies selective sharing across the most area-critical components, including a shared 25x64-bit register bank, shared...
Note on Number-Theoretic Transforms for Implementers -- Butterflies, Twisting, Incompleteness, and Good's Trick
Bo-Yin Yang
Implementation
We develop (mostly) the radix-2 number-theoretic transform (NTT) and
its butterflies, the twisting trick and why it never changes the
transform, the freedom to use Cooley--Tukey butterflies in both
directions, incomplete NTTs, Good's trick, and the ways all of these
combine---closing with the coefficient-bound bookkeeping that
motivates the whole toolkit. This note is intended to help
implementers of postquantum cryptography, and is compressed from the
author's lecture...
Falcon Verify on AVX-512: Speed Records
David Rubin, Emanuele Cesena
Implementation
We present a fast implementation of Falcon (FN-DSA) signature verification with AVX-512. On a modern AMD Zen5 core, it completes a Falcon-512 verification in 3.6 microseconds, 2.6 times faster than an already optimized baseline, with comparable gains on Zen4, and consistent results across clang 21 and gcc 15.
The speedup comes from rewriting the Number-Theoretic Transform (NTT) and from vectorising all other stages of the verification algorithm. The novelty is to use a 32-bit...
Toward a Secure Fixed-Point Implementation of the Falcon Signature Scheme
Daniel De Almeida Braga, Pierre-Alain Fouque, Bachir Lachguel, Thomas Prest
Implementation
Falcon was selected by NIST in 2022 for standardization as a post-quantum digital signature scheme. Among all standardized signature schemes, Falcon achieves the smallest signature size. Its main drawback, however, is its reliance on floating-point arithmetic, which plays a critical role in the security analysis. This reliance poses significant challenges for practical implementations: some platforms lack floating-point units, floating-point division is not constant time on many processors,...
BF²: A Bloom-Filtered Brute-Force Framework for Multi-Target Password Recovery
Cansu Karakuzu Aslan, Wenzel Pünter, Christian Dörr
Implementation
Password-based authentication remains widespread, and large-scale sets of leaked hashes enable practical offline brute-force attacks. Multi-target attacks, which check candidates against large sets of hashes simultaneously, are particularly effective. Understanding the capabilities of low-cost platforms for such attacks is important to assess real-world password security risks.
Therefore, we present BF², a modular and scalable FPGA–CPU framework that accelerates multi-target password...
Oblivious Sorting under Fully Homomorphic Encryption: A Comprehensive Survey and Performance Analysis
Omar Ahmed, Rostin Shokri, Nektarios Georgios Tsoutsos
Implementation
Outsourcing computations to cloud providers raises significant data privacy concerns, making Privacy-Preserving Computation via Fully Homomorphic Encryption (FHE) increasingly vital. However, adapting data sorting routines to the FHE domain introduces severe performance bottlenecks. This survey systematizes the state-of-the-art in FHE-based sorting algorithms. A novel complexity metric, FHE-Effort, is introduced to accurately evaluate homomorphic circuit efficiency. Eighteen algorithms are...
On $k$-way split multiplication algorithms
Mehmet Özgün Cihangir, Oğuz Yayla
Implementation
Efficient polynomial multiplication and matrix-vector operations are fundamental to computational algebra and modern cryptography. In lattice-based post-quantum cryptography (PQC), schemes utilizing Number Theoretic Transform (NTT)-unfriendly rings require highly optimized subquadratic multiplication algorithms. In this paper, we establish a rigorous mathematical framework for generalized $k$-way split polynomial multiplication and Toeplitz Matrix-Vector Product (TMVP) algorithms over...
Exploiting Load/Store Leakage of Sparse Vectors for Key Recovery in HQC
Gustavo Banegas, Benjamin Smith, Jad Zahreddine
Implementation
Hamming Quasi-Cyclic (HQC) is a code-based key encapsulation mechanism
selected by NIST for standardization,
making its resistance to implementation attacks critically important.
We present a side-channel attack that exploits load/store leakage
in the manipulation of HQC's sparse secret vectors.
Analysing Cortex-M4 assembly generated from the reference
implementation, we identify a leakage surface in which the low and
high 32-bit halves of each 64-bit word...
From PQC to HHE: Reusing a Co-Design Platform for Side-Channel-Protected PASTA
Ahmet Malal, Tolun Tosun, Oğuz Yayla, Erkay Savas
Implementation
Hybrid homomorphic encryption (HHE) lets a constrained client send compact symmetric ciphertexts while a server transciphers them into homomorphic ciphertexts, making HE-friendly ciphers such as PASTA a practical choice. Efficient and side-channel-secure execution of PASTA on embedded devices, however, remains challenging, since existing hardware relies on dedicated cipher cores and provides no side-channel protection. We present a hardware/software co-design of PASTA on RISQrypt, an...
MQ on my Hardware: Performance Analysis of MQOM on FPGA
Stelios Manasidis, Quinten Norga, Suparna Kundu, Ingrid Verbauwhede
Implementation
Recent algorithmic advancements in the Multi-Party Computation-in-the-Head (MPCitH) paradigm have resulted in more efficient post-quantum digital signature schemes. MQOM is a MPCitH-based digital signature scheme and candidate in the ongoing NIST Post-Quantum Cryptography (PQC) standardization effort, offering performance competitive with lattice- and multivariate-based schemes in software.
In this work, we develop a dedicated hardware accelerator for MQOM and analyze the impact of recent...
88-XOR Implementation of the AES MixColumns Matrix
Jérémy Jean
Implementation
We give in this short note a circuit implementing the matrix-vector product with the 32x32 binary matrix of the AES MixColumns using 88 XOR gates. Previously known circuits minimizing this metric have been published in the past years and achieved 94 XOR, 92 XOR, 91 XOR, and 89 XOR. As far as we can tell, a circuit with 88 XOR was previously unknown.
Scalable High-Throughput FPGA Architecture for SMAC Message Authentication Code
Ahmet MALAL, Hakan Güler, Bahadır Aydoğan, Oğuz Yayla
Implementation
SMAC is a recently proposed by Wang et al.~stand-alone Message Authentication Code (MAC) constructed from repeated applications of the AES round function and featuring an aggregation mode, SMAC-1$\times n$, for scalable parallel processing. Although originally designed for high-throughput CPU implementations leveraging AES-NI instructions, its structural properties suggest strong compatibility with hardware parallelism. However, no systematic FPGA-oriented architectural study of SMAC has...
Hybrid hash function based on the DLP and SIS problems
Dimitri Koshelev, Francesc Sebé
Implementation
This short note discusses in detail a folklore but little-known hybrid hash function grounded on both the discrete logarithm and short integer solution problems. In particular, specific satisfactory parameters are provided to ensure the standard $128$-bit security level for the lattice problem with $256$-bit module, which may be useful in its own right. The hash function is a natural generalization of the classical Pedersen and Ajtai ones. Nevertheless, to the authors' knowledge, no one has...
A High-Speed Hardware Accelerator for QR-UOV Signature Scheme
Renma Sugai, Hiroshi Amagasa, Rei Ueno, Naofumi Homma
Implementation
This paper proposes a high-speed hardware accelerator for QR-UOV, a multivariate scheme, that executes all three operations: key generation, signature generation, and signature verification.
QR-UOV utilizes a quotient polynomial ring structure to reduce the public-key size of the original UOV scheme; however, this introduces functional requirements distinct from other multivariate schemes, such as polynomial-matrix operations over $\mathbb{F}_{q^\ell}$, coefficient expansion for the...
Lightweight Hardware Accelerator for the UOV Signature Scheme with Oil Space Blinding
Florian Krieger, Maciej Czuprynko, Sujoy Sinha Roy
Implementation
In reaction to the emerging quantum threat, the National Institute of Standards and Technology (NIST) seeks post-quantum secure digital signature schemes. NIST's ongoing competition recently advanced to the third round, in which the Unbalanced Oil and Vinegar scheme (UOV) is a promising candidate due to UOV's conservative design, small signatures, and performant signing and verification. While these benefits make UOV attractive, the implementation aspects for compact hardware acceleration of...
Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4
Jihoon Jang, Hanbeom Shin, Suhri Kim, Seokhie Hong, Donggeun Kwon
Implementation
In this paper, we present an optimized implementation of Hamming Quasi-Cyclic (HQC) on the ARM Cortex-M4. We optimize (i) the polynomial multiplication and (ii) the support expansion in fixed-weight sampling, and (iii) propose an optional caching strategy that reuses the public transforms and hash recomputed under a fixed key.
For the polynomial multiplication, the fixed-constant multiplications in the Frobenius additive FFT (FAFFT) butterfly spend nearly half of their instructions on VMOV...
Quantum Circuit Optimization with LLMs under a Structured Guideline
Kyungbae Jang, Hyunji Kim, Hwajeong Seo, Anupam Chattopadhyay
Implementation
The cost of quantum cryptanalysis is dominated by the quantum circuit of the target cipher. Estimating the quantum attack cost of a cipher thus requires building that circuit and measuring its qubit count, Toffoli count, and Toffoli depth. This is manual work that needs expert knowledge and must be redone for each cipher and each cost target. Large language models handle ordinary programming well, but their use in constructing quantum circuits for ciphers is still limited. In this work, we...
Coupling Leakage in Theory and Practice - Unveiling (Post-PnR) Security Flaws in Masked FPGA-Mapped Designs
Nicolai Müller, Daniel Lammers, Simon Osterheider, Amir Moradi
Implementation
With the widespread adoption of Field Programmable Gate Arrays (FPGAs) in security-critical industries such as defense and telecommunications, ensuring the confidentiality of sensitive data processed by these devices has become paramount. Side-Channel Analysis (SCA) poses a significant threat, necessitating the protection of cryptographic primitives through effective and efficient countermeasures. Within the framework of well-established formal adversary models, Boolean masking offers...
CMALU: Compact Fault-Tolerant Modular Arithmetic Logic Unit for Post-Quantum Cryptography
YoungBeom Kim, Malik Imran, Zain Ul Abideen, Ciara Rafferty, Ayesha Khalid, Máire O’Neill, Seog Chung Seo
Implementation
The rise of quantum computing threatens widely deployed public-key cryptosystems, driving the adoption of post-quantum cryptography (PQC) algorithms that rely heavily on modular arithmetic. Existing hardware accelerators of the PQC algorithms for resource-constrained Internet-of-Things (IoT) devices remain limited and lack integrated fault detection mechanisms. In this work, we present CMALU, a Compact, fault-tolerant Modular Arithmetic Logic Unit supporting six operations on a single...
LESS on the Cortex-M4: Characterizing the Speed–Memory Design Space of Code-Equivalence Signatures
Minwoo Lee, Minjoo Sim, Subeen Cho, Yulim Hyoung, Hwajeong Seo
Implementation
LESS is a code-based signature scheme built on the linear equivalence problem and, in its v2.0 round-2 form, a candidate in the NIST call for additional post-quantum signatures. No microcontroller implementation of it has been reported: the official benchmarking effort for the additional signatures excluded it on memory grounds, and an x86-massif cross-check puts the reference's peak stack at up to $\approx 836$~KB---beyond the SRAM of even the largest mainstream Cortex-M4. This paper...
ML-QED-Lite: A Lightweight Machine Learning-Based Tool for Supporting Post-Quantum Cryptography Migration in Executable Binaries
Seung-Won Lee, Hwa-Jeong Seo
Implementation
To initiate migration to post-quantum cryptography (PQC), it is necessary to identify whether deployed software uses quantum-vulnerable (QV) public-key cryptographic schemes such as RSA, ECDSA, and Diffie–Hellman (DH). However, many ELF executables are distributed without source code, making it necessary to directly screen executable binaries for QV candidates. A prior tool, Quantum-vulnerable Executable Detection (QED), provides high precision but incurs substantial analysis cost, whereas...
CT-KAT: A Multilayer Analysis Platform for Automated Screening of Constant-Time Risks in PQC C Implementations
Seung-Won Lee, Min-Seo Kim, Su-Min Jeong, Hwa-Jeong Seo
Implementation
Following the standardization of major post-quantum cryptography (PQC) algorithms, C implementations of ML-KEM, ML-DSA, and SLH-DSA have been rapidly deployed. However, known-answer tests (KATs) verify only functional correctness and do not establish the absence of timing leakage caused by secret-dependent branches, memory accesses, or variable-latency instructions. This paper presents CT-KAT, an integrated screening platform for assessing constant-time risks in PQC C implementations. CT-KAT...
Accelerating the AIMer Post-Quantum Signature with AVX-512: A Field–Keccak Speedup Analysis
Seung-Won Lee, Si-Woo Eum, Hwa-Jeong Seo
Implementation
AIMer is a post-quantum digital signature scheme with a conservative design. Its security relies only on the symmetric-key one-way function AIM2 and an MPC-in-the-Head (MPCitH) zero-knowledge proof. AIMer is a Korean post-quantum cryptography (KpqC) standard. However, the AIMer standard code released in January 2026 is a portable C reference implementation. It does not include processor-specific optimizations. As a result, it does not exploit AVX-512, a 512-bit vector instruction set...
Beyond Size: Do Hybrid PQC Certificates Actually Enforce the Classical–PQC Binding? A Cost-and-Security Study
Minwoo Lee, Minjoo Sim, Siwoo Eum, Subeen Cho, Yulim Hyoung, Hwajeong Seo
Implementation
As TLS 1.3 migrates to post-quantum cryptography (PQC), hybrid X.509 transition strategies—alternative-signature (Catalyst), Composite, Chameleon, and signature combiners—are compared on cost but rarely on whether they actually enforce the classical↔PQC binding they promise. We show they often do not, and that the failure persists even in stacks that do check the binding. The same BouncyCastle library accepts a Catalyst certificate carrying a forged ML-DSA signature on its default path yet...
Optimized Implementation of Warp-Cooperative GPU HCTR2-ARIA Wide-Block Encryption
Siwoo Eum, Minho Song, Seung-Won Lee, Hagyeong Kim, Hwajeong Seo
Implementation
HCTR2 is a wide-block encryption mode that encrypts one fixed-size message as a single unit, so that flipping a single plaintext bit re-randomizes the whole ciphertext. Its main use is disk encryption, where the message is a disk sector. We instantiate it with ARIA, the Korean national block-cipher standard, and implement it on an NVIDIA RTX 4080 GPU. With many independent messages, assigning one thread per message keeps the device occupied. At low queue depth, however, most of the GPU sits...
Evaluating Hybrid KEM/DSA for KpqC and NIST PQC on ARM Cortex-M4
Minjoo Sim, Minwoo Lee, Subeen Cho, Yulim Hyoung, Hwajeong Seo
Implementation
Primitive-only PQC benchmarks are insufficient for attributing composed hybrid costs on Cortex-M4 because shared hash backends, randomized-signature behavior, and fixed classical/wrapper work affect measured performance. We implement a common bare-metal Cortex-M4 harness for representative KpqC/NIST families, measuring uniform Hash-CT hybrid KEM benchmark rows with X25519 and Bindel et al. hybrid-signature AND-combiner rows. The goal is composed-cost attribution under a uniform benchmark...
Optimizing ARIA-GCM on GPUs
Min-Ho Song, Si-Woo Eum, Seung-Won Lee, Ha-Gyeong Kim, Hwa-Jeong Seo
Implementation
This paper proposes an optimized GPU implementation of the ARIA-GCM authenticated-encryption pipeline (CTR keystream, GHASH authentication, and their AEAD composition): ARIA-CTR uses a packed 32-bit S-box staged in shared memory, GHASH is optimized separately with a fixed-key 4-bit Shoup lookup table, the two stages are integrated as both a two-kernel and a fused single-kernel AEAD, and the same aria_gcm.cu source is tuned for Ampere and Pascal through compile-time parameters. For ARIA-CTR,...
Quantum Implementation and Analysis of Rijndael
Gyeongju Song, Hwajeong Seo
Implementation
We present a quantum resource estimation of the Rijndael variants
\[
N_b = N_k \in \{4,5,6,7,8\}, \qquad N_r = N_b + 6,
\]
under the NIST MAXDEPTH quantum cost model. Extending the AES quantum
encryption oracle~\cite{ref5} parametrically to arbitrary $N_b = N_k$,
we generalize the in-place key schedule, including the single- and
double-\texttt{SubWord} cases, the \texttt{ShiftRows} offsets, and the
round constants. We implement and verify the resulting oracles using
ProjectQ. The...
Improved Quantum Circuits for Information Set Decoding with Application to Code-Based Cryptography
Hyunji Kim, Kyungbae Jang, Hwajeong Seo
Implementation
Information set decoding (ISD) is the standard generic decoding attack considered for code-based cryptography. A concrete quantum-resource estimate for Grover-accelerated ISD requires an oracle whose dominant component is Gauss–Jordan elimination.
We improve the elimination circuit of Perriello et al. [25] and Jang et al. [15] by not updating the entries that no later pivot or the final weight predicate reads. The required result vector is recovered by a parallel back-substitution on the...
A Memory-Efficient and Assembly-Optimized Implementation of NTRU+
SuBeen Cho, Jiwon Bang, Minjoo Sim, Hwajeong Seo
Implementation
This paper presents a memory-efficient and high-speed implementation of NTRU+, one of the key encapsulation mechanisms (KEMs) selected by Korea’s post-quantum cryptography project (KpqC), on the ARM Cortex-M4. NTRU+ is small enough to run on its own on a Cortex-M4 class microcontroller, yet in real embedded environments, the peak stack occupied by polynomial buffers and the running time dominated by the NTT become key constraints. To address this, in the proposed technique, we reduce memory...
Accelerating FAEST Signing on GPU via Fused AES Constraint Generation and Batched Leaf Hashing
Ha-Gyeong Kim, Si-Woo Eum, Seung-Won Lee, Ui-Jae Kim, Min-Ho Song, Hwa-Jeong Seo
Implementation
FAEST is a symmetric-key post-quantum digital signature scheme and a third-round candidate in the NIST Additional Digital Signatures standardization process. Its signing path concentrates cost in two operations: round-wise constraint generation, which proves in zero knowledge that the AES circuit is computed correctly, and finite-field multiplication, which computes the leaf nodes of a vector commitment. This paper accelerates both operations on a CUDA-enabled GPU, with AES round constraint...
Reliable TRNG and its Challenges
Raja Adhithan Radhakrishnan
Implementation
The objective of this work is to investigate methods
for improving the self-tuning mechanism of ring oscillator (RO)
based True Random Number Generators (TRNGs). It also
examines the challenges involved in achieving a reliable and
stable design over long-term operation. Furthermore, this work
analyzes potential approaches to address these challenges and
validates their effectiveness using the NIST statistical test suite.
Chimera: A Hybrid GPU Backend for Sumcheck Acceleration in Zero Knowledge Provers
Kashfia Farheen, Nektarios Georgios Tsoutsos
Implementation
Zero-knowledge proof systems are increasingly relying on the Sumcheck protocol to avoid the FFT-heavy structure of earlier SNARK designs. Sumcheck is well suited for GPU acceleration; it consists of sequential rounds where each round performs regular, parallelizable operations over large multilinear evaluation tables. The focus is on how to organize this work across rounds: intuitively, the active polynomial state should remain close to the device that processes it, the CPU-GPU boundary...
Reducing Multiplicative Complexity via Conjugate Cipher
Noémie Akpaki, Nicolas DAVID
Implementation
Multiplicative complexity have shown to be an important metric for efficient implementations in various contexts such as side-channel secure implementation and transciphering.
We introduce a generic framework based on conjugacy to reduce the multiplicative complexity of block ciphers. Our approach exploits the iterative structure of the block cipher to build alternative implementation based on conjugate round operations with overall smaller multiplicative complexity.
We apply this...
Slicing Bits and Cutting Costs in CDT Sampling: High-Order Masking of FrodoKEM's Gaussian Sampler, Revisited
Calvin Abou Haidar, Thomas Espitau, Clément Hoffmann, Mehdi Tibouchi
Implementation
FrodoKEM, a key encapsulation mechanism based on the standard (unstructured) LWE assumption, is recommended as a conservative choice for post-quantum key exchange by agencies like BSI and ANSSI. As such, it has garnered substantial attention from an implementation security standpoint. In particular, several papers have looked into masking FrodoKEM, and, like for various other lattice-based cryptosystems, identified the Gaussian sampling operation as a major bottleneck. In FrodoKEM, it is...
A Prototype-Based Study of Zero-Knowledge Proof Verification for Privacy-Preserving Blockchain Interoperability
Chilume O. Gabriel, Hlomani B. Hlomani, Kabo Nkabiti
Implementation
Blockchain networks need to exchange messages and assets across independent systems, but cross-chain verification can expose private validation data to relayers, bridge logic, validators, or destination-chain components. This paper presents a prototype-based zero-knowledge verification layer for privacy-preserving blockchain interoperability. The prototype uses Circom and SnarkJS to generate Groth16 proofs, verifies those proofs in Rust using arkworks BN254, and maps the result into a...
Optimization of Hardware Architecture for Quantum Key Distribution
Raja Adhithan Radhakrishnan
Implementation
The main objective of this paper is to acceler
ate the post-processing of Quantum Key Distribution (QKD)
using an energy-efficient pipelined architecture implemented
on a Field-Programmable Gate Array (FPGA). The proposed
architecture aims to improve processing speed while efficiently
utilizing hardware resources. In addition, this work compares
the proposed approach with existing approaches to demonstrate
its performance and resource efficiency.
Efficient multiplication over $\mathbb{F}_2[x]$ is a core primitive in classical and post-quantum cryptographic software. As ARM and RISC-V become increasingly relevant for open-source cryptographic libraries such as OpenSSL and liboqs, arithmetic kernels must be retuned across a wider range of Instruction Set Architectures (ISAs). High-performance arithmetic libraries recursively apply Karatsuba- and Toom-style decomposition rules, each splitting the operands into smaller subproblems and...
Established multiple-precision software generally reserves Fast-Fourier Transform (FFT) multiplication for very large operands; for example, GMP reports full-product FFT thresholds of roughly 3000--10000 limbs. Becker et al. showed that Number-Theoretic Transform (NTT) multiplication can become advantageous much earlier on Cortex-M microcontrollers. It thus becomes an interesting engineering question to check where the crossover actually takes place on a big modern CPU. Our first-generation...
The practical performance of Gröbner basis implementations depends on the polynomial systems being solved. This work evaluates 20 open-source imple- mentations with Magma as a proprietary reference on systems derived from small-scale AES over GF(2). The Easy, Medium, and Hard experiments use field equations and ten input instances each, comparing wall-clock runtime, peak memory, success rate, and CPU usage. Implementations with no successful runs are excluded from subsequent...
The Hamming-Quasi Cyclic (HQC) is the Post-Quantum Cryptography (PQC) Key Encapsulation Mechanism (KEM) algorithm recently selected by the National Institute of Standards and Technology (NIST) for standardization as an alternative candidate for the KEM category. While the algorithm is mathematically robust, several studies state that the physical implementation of HQC remains vulnerable to power-based side-channel attacks (SCA). Countermeasures for such SCAs primarily focus on masking and...
Bootstrapping is the main performance bottleneck in bitwise Fully Homomorphic Encryption (FHE), and practical acceleration requires careful orchestration of the blind rotation and external product chain under GPU resource constraints. This paper presents BOLT-FHE, a GPU bootstrapping framework that emphasizes block-local execution, on-chip tiling, and a unified MegaKernel supporting both gadget decomposition and modulus raising, with optional support for a recently proposed technique...
Applying a GPU to FAEST v3 signing requires more than exploiting parallel computation: challenge grinding must preserve the minimum accepting counter selected by the original signing procedure, and single-sign latency and multi-sign throughput must be evaluated under distinct execution conditions. We implement GPU-assisted challenge grinding in which the GPU filters candidates over contiguous counter ranges in parallel, while the CPU performs the final acceptance checks in increasing counter...
X-Wing combines ML-KEM-768 and X25519 for hybrid key establishment. On Cortex-M85, we adapted the coefficient representation and storage order of an MVE NTT to the existing ML-KEM path and optimized data-format transformations and X25519 arithmetic while preserving draft-10. During encapsulation, the two X25519 results share an inversion. We compared the eight-technique integrated implementation B8 with the baseline A0 in the same executable. Across ten paired comparisons from five...
With FIPS 205 standardizing post-quantum signatures and ARMv8 spanning mobile, embedded, and server platforms, fast hash-based signing on ARM has become a pressing need. Since these schemes are dominated by hash calls, hash optimization largely determines performance. K-SPHINCS+ replaces the internal hash of stateless SPHINCS+ with LSH, a Korean standard hash family. However, prior K-SPHINCS+ work provides only reference C code, and existing ARMv8 LSH implementations mainly optimize...
MAYO is one of nine third-round candidates in NIST’s process for additional post-quantum signatures. Its third-round specification of August 2026 changed the level-1 parameters and added a hash-derived linear term Λ to the verification equation, so earlier optimised implementations no longer match it. We present an optimised NEON implementation of MAYO Round 3 for the Apple M1 and compare it with the upstream NEON code on the same machine. A generator emits the sixteen GF(16) kernels of each...
q0 preprocessing verification and native HE computation verification can each succeed without establishing that their results refer to the same canonical source. To address this source mismatch across verification paths, we propose a verifier-aligned public-input transfer procedure that links heterogeneous authenticated artifacts to a single canonical-source relation. The public verifier reconstructs challenges and targets from the actual statements, roots, approved verification keys, and...
The quantum threat to traditional public-key cryptography makes migration to post-quantum cryptography necessary. During migration, the presence of both post-quantum and traditional (PQ/T) algorithms does not establish that their outputs jointly determine a key or verification decision. We propose HyBind, an intraprocedural Python source-analysis framework using registered API contracts. HyBind traces component secrets to returned keys and checks whether signature acceptance requires both...
Migration to post-quantum cryptography requires knowing not only which cryptographic libraries are present on a system, but which algorithms are actually executed at runtime. Static binary analysis answers the former question; the latter requires dynamic observation. We present a runtime cryptographic discovery system for Linux, built on eBPF user-space probes, that detects classical and post-quantum operations as they execute. Because key generation and encapsulation consume fresh...
Certificate issuance and per-connection signing impose different costs in TLS 1.3, and authentication cost can change depending on how signature algorithms are assigned to Root, Intermediate, and Leaf roles. PQC migration should therefore consider not only algorithm performance, but also role placement, certificate structure, network conditions, and server load. We evaluate all 13³ role placements formed by 13 signature parameter sets while holding the algorithm composition fixed and...
The command and authentication channels of unmanned systems (UAVs and UGVs) still rely on classical public-key cryptography, which large-scale quantum computers threaten. Migrating these channels to post-quantum cryptography (PQC) is therefore imperative, yet PQC keys and signatures are substantially larger, and sometimes slower, than their classical counterparts, while command links are tightly constrained in message rate, maximum transmission unit (MTU), latency deadline, and bandwidth....
Post-quantum migration is no longer only an algorithm problem but a deployment problem: standardized schemes must be dropped into real protocol stacks, on real (often resource-constrained) targets, without giving up memory safety. The Korean Post-Quantum Cryptography (KpqC) competition standardized two KEMs (NTRU+, SMAUG-T) and two signature schemes (HAETAE, AIMer), distributed as memory-unsafe C. Building on our native-Rust KpqC implementation, we expose the schemes through a single...
Hybrid key exchange and Composite ML-DSA signatures in TLS 1.3 split one operation into a classical component and a post-quantum cryptography (PQC) component that do not depend on each other. Using a cryptographic workload of the TLS handshake rather than a full TLS connection, this paper measures whether running the two components at the same time on two pthread worker threads reduces latency compared with sequential execution. We evaluate the three ECDHE–ML-KEM groups of RFC 10024 with the...
Cryptographic tokens that retain RSA credentials for compatibility with existing authentication systems require efficient private-key operations on small processors. Cortex-M0+ (M0+) lacks native widening multiplication, whereas Cortex-M3 (M3) provides it with operand-dependent latency. We optimize Montgomery multiplication and squaring using 15-bit limbs and integrate the kernels into RSA-2048 and RSA-3072 private operations based on the Chinese remainder theorem (CRT). On M0+, a wrapped...
End-to-end cycle counts quantify lattice key-generation time on a microcontroller, but not which implementation decisions create that cost. We develop a controlled attribution method for public structure, candidate admission, and transform lifetime: how public data are organised, when candidate acceptance is checked, and whether transformed secret state is retained or reconstructed. On a fixed STM32L476RG Cortex-M4 target, paired runs within one executable keep the relevant secret, accepted...
ECLIPSE is the prime-degree signature construction that the PRISM authors describe next to their salted scheme, with that salt, implemented. It exists because a two-dimensional isogeny signature carries an auxiliary isogeny that the hash does not bind: the SQIsign round-3 specification states that SQIsign cannot achieve strong unforgeability for this reason, and PRISM's prime-degree variant was designed without the auxiliary isogeny but published without parameters, code or a verifier that...
Programmable bootstrapping is a central procedure in the FHEW and TFHE families of fully homomorphic encryption schemes, but its computational cost remains a major performance bottleneck. Multi-value bootstrapping amortizes this cost by sharing a blind rotation among several functions evaluated on the same encrypted input. However, the standard choice of common factor for multi-value bootstrapping can produce unnecessarily large noise amplification for many collections of lookup tables,...
We present the first formal verification results for the highly optimized additive Fast Fourier Transforms (FFTs) in two different code-based post-quantum cryptosystems: Classic McEliece and Hamming Quasi-Cyclic. Since Classic McEliece is already a standard and HQC is being standardized, they may both see wide adoption in the future. The Gao-Mateer and Frobenius additive FFTs during decoding are clearly the most intricate, error-prone components in the two cryptosystems respectively. They...
Discrete Gaussian sampling remains one of the most delicate operations to protect against side-channel attacks in lattice-based cryptography. In this work, we present the first masked evaluation of the Knuth-Yao sampler, a generic building block for lattice-based schemes. Its random-walk formulation might seem fundamentally at odds with masking, since the walk's control flow depends precisely on the secret sample being produced. We show instead that its underlying tree structure is...
CROSS is a code-based signature scheme built on the Restricted Syndrome Decoding Problem and is a second-round candidate in NIST's additional digital signature standardization process. Recent work shows that its reference implementation is vulnerable to side-channel analysis, allowing an attacker to recover the long-term secret key from a single power trace. No masking countermeasure has so far been designed for CROSS's restricted-syndrome framework, leaving its practical side-channel...
We present SUCRE, a novel countermeasure designed to physically protect the rejection sampling step of ML-DSA, one of the post-quantum signature schemes standardized by NIST. At the core of SUCRE is a masking gadget that securely unmasks a vector while simultaneously applying a random permutation of its coefficients. This lightweight mechanism preserves the vector’s infinity norm, enabling rejection sampling to proceed as usual without requiring any complex mask conversions. We formally...
SQIsign holds the smallest known combination among post-quantum signatures, and its dimension-2 signature carries an auxiliary curve whose only role is to make the response computable in dimension 2. SQIsignHD's dimension-4 response has no such curve. We take that construction to the round-3 SQIsign primes, which replaced the round-2 primes in September 2026 after the attacks of Wesolowski and others, on our Rust port of the SQIsign reference code, and measure it in one session next to the...
SQIsign samples a random ideal of odd prime norm at every key generation and every signature through $\mathbf{PrimeNormIdealInertSampling}$. We revisit this step from a new perspective: the inert ideals of a given norm form a circle modulo the norm, and drawing lines through a fixed point of this circle turns a single random parameter into an ideal, hitting each ideal exactly once. Compared with the reference implementation, the new algorithm requires one modular inversion and a handful...
AIMer v3 replaces AIM2 with AIM3. We test whether SIMD methods developed for AIM2 can accelerate AIM3 while preserving its results. We port the field-arithmetic, affine, party-parallel, and SHAKE methods of the earlier AVX-512 implementation to AIMer v3 and implement the same operations for AVX2. For Cortex-M55 MVE, we use 16-bit polynomial partial products for field multiplication, group corresponding 32-bit words from four parties, and vectorize both affine layers, including the input...
Efficient execution of Post-Quantum Cryptography (PQC) in resource-constrained embedded environments requires the joint optimization of not only algorithm vectorization but also dataflow and instruction execution order. This paper implements the HAETAE-2/3/5 and SMAUG-T1/3/5 parameter sets of two Korean Post-Quantum Cryptography Competition (KpqC) winners using the M-profile Vector Extension (MVE) of the Arm Cortex-M55 and analyzes the contribution of each optimization stage to performance...
Many isogeny-based schemes rely on the Deuring correspondance and thus require computations with quaternion ideals. In this paper, we generalize the recent results of Leroux on the inert representation of quaternion ideals, yielding a normalized representation together with a set of simple and efficient algorithms for quaternion ideals in the context of the Deuring correspondence when the prime characteristic $p$ is equal to $3 \bmod 4$. One of the main benefit of our new algorithms...
Key generation in ML-KEM (CRYSTALS-Kyber) samples a short secret from a centered binomial distribution (CBD) and immediately transforms it with the number-theoretic transform (NTT). Each execution draws fresh randomness, so an attacker obtains a single trace and cannot average. We show that one power trace of the optimized pqm4 implementation on an Arm Cortex-M4 suffices, and that the two operations leak complementary information. The CBD sampler stores each coefficient as a signed $16$-bit...
Optimizing quantum circuits is critical: circuits must fit within the resource limits of a quantum computer, and every unnecessary operation increases their cost and probability of failure.We present a simple optimization algorithm for quantum circuits that (1) is very fast, (2) scales to millions of operations, and (3) matches or outperforms the optimization quality of the best existing optimizers and superoptimizers. Our key insight is that we can compactly represent circuit equivalence...
Binomial coefficients modulo an integer, $\binom{N}{R} \pmod m$, are a primitive of combinatorial counting, yet the two textbook methods collapse at scale: the Pascal recurrence costs $\Theta(NR)$ time, and the factorial-table method costs $\Theta(N)$ memory, requires a prime modulus, and requires $N < m$. Composite moduli are harder still, because factorials are not invertible modulo prime powers and the exact power of $p$ dividing the coefficient must be tracked. We present a complete,...
In this work, we present a quantum implementation of Keccak-$f(25)$, a toy version of SHA-3 introduced in the specification of Keccak, with the objective of using as few qubits as possible so that the resulting implementation can be run both on quantum emulators and on quantum hardware. The code, written in Qiskit, uses $25$ qubits representing the internal state. All computations are performed in-place, meaning that no ancillary qubit is required. We use this implementation to...
The practicality of post-quantum cryptography (PQC) on IoT devices depends on both operation cycle counts and communication costs from public values (public keys, ciphertexts, and signatures). The DSA HAETAE and the KEM SMAUG-T, both selected in the KpqC competition, provide smaller public values than ML-DSA and ML-KEM; however, the lack of platform-specific optimization leaves their operation cycle counts high, which can offset this advantage. In this paper, we optimize HAETAE and SMAUG-T...
Streamlined NTRU Prime (sntrup761) is a lattice-based key encapsulation mechanism that, although not a NIST standard, remains widely deployed in critical internet infrastructure. It is the post-quantum key-exchange default in OpenSSH, standardized in RFC 9941, and used well beyond SSH, in Red Hat Enterprise Linux, the liboqs library, PQConnect, and commercial VPNs. Its decapsulation performs a polynomial multiplication over the characteristic-three ring $\mathbb{Z}_3[x]/(x^{p}-x-1)$, so...
Evaluation keys represent a primary memory and initialization bottleneck in matrix-native fully homomorphic encryption (FHE). In the Gentry–Lee (GL) framework, each Trace product yields a four component ciphertext whose BigSwitch procedure requires two extended-ring keys, dominated by a massive product-secret key (sXsY → sX). We present Trace-Factored BigSwitch (TFB), which structurally eliminates this product-secret evaluation key by exploiting the rank one tensor structure of the...
Falcon is one of the 3 post-quantum signature schemes already selected by NIST for standardization (as FN-DSA). It is very compact and efficient, but also infamously difficult to implement correctly and securely. This is due in particular to its reliance of various floating point operations, the most complex and costly of which are square root computations. In this paper, we first point out that those square root computations are in fact wholly unnecessary: the algorithm can be rewritten...
Masking is a standard countermeasure against side-channel attacks on embedded cryptographic implementations. Its security is commonly analyzed in the random probing model, which offers a useful trade-off between realistic leakage assumptions and tractable security proofs. Recent years have seen the emergence of several masking compilers based on compositional security frameworks such as general/cardinal random probing composability (RPC). Most of these approaches rely on dedicated refresh...
Verifiable oblivious pseudorandom functions (VOPRFs) enable a client to evaluate a pseudorandom function on a private input under a server-held key while verifying that the server evaluated the function honestly using its committed key. Existing practical VOPRFs rely mainly on assumptions that are vulnerable to quantum attacks. Prior lattice-based constructions achieving both round optimality and malicious security have remained theoretical proposals without concrete implementations, largely...
Side-channel leakage certification aims to quantify what an attacker can learn about a secret variable from observed leakage. Existing information-theoretic estimators, such as perceived information (PI), hypothetical information (HI), and nonparametric mutual information (MI), aim to quantify distributional leakage, but they become unstable in high-dimensional traces and do not provide a finite-sample certificate of the best attacker. We introduce \emph{Bounded Information} (BI), a probably...
We present a RAM-efficient implementation of Falcon: RAM usage has shrunk to about 11 kB, down from about 31 kB in the previous implementation of Falcon-512. This code is furthermore faster on Arm Cortex M4, with average signature generation cost down to 13.45 million cycles. Optimization techniques include a novel variant of the FFT, replacement of some floating-point operations with modular integer computations, delayed addition of input within the Fast Fourier sampling process, and an...
We develop the method of Luo, Fu and Gong (LFG) - as extended by Fan, Kuchta, Sica and Xu (FKSX) to use endomorphism scalars - in order to find best families suitable for multi-scalar multiplication (MSM) for any number $n$ of points and with adjustable storage. In particular we lower storage requirements by an average of 67% and decrease the number of curve operations by an average of 7% (and up to 10.6%), relative to the LFG method. Compared to Pippenger's variant (standard when $n$ is...
We consider methods for scalar multiplication on an elliptic curve where the scalar digits are processed from left to right, that is, from most significant to least significant. We analyze exceptions that may arise during point addition and doubling throughout the multiplication. Eliminating such exceptions is critical for achieving constant-time execution and preventing timing attacks. We establish conditions under which exceptions occur only at the final addition or not at all. Guided by...
Cloud computing has been a prominent technology that allows users to store their data and outsource intensive computations. However, users of cloud services are also concerned about protecting the confidentiality of their data against attacks that can leak sensitive information. Although traditional cryptography can be used to protect static data or data being transmitted over a network, it does not support processing of encrypted data. Homomorphic encryption can be used to allow processing...
Fully homomorphic encryption (FHE) allows a server to run inference directly on encrypted data, making it a promising foundation for private transformer inference. Its dominant scheme, CKKS, has no native matrix multiplication, so the encrypted matrix multiplications at the heart of transformers dominate inference cost. The recent GL scheme supports matrix multiplication natively, but its ciphertexts, plaintexts, and evaluation keys are so large that a direct GPU implementation would require...
Under-constrained arithmetic circuits are a recurring source of soundness failures in zero-knowledge applications: after fixing the public statement, a malicious prover may be able to assign a security-relevant wire in more than one way while still satisfying the circuit. Existing tools attack this uniqueness question with solver-based checking, direct polynomial solving, abstract interpretation, or fuzzing. We study a complementary algebraic diagnostic based on exact Jacobian linear...
Mutual TLS (mTLS) authenticates both peers and therefore incurs post-quantum signature costs on every connection. Concurrent handshakes expose independent ML-DSA operations, but executing them jointly is difficult: signing is rejection-divergent, verification uses heterogeneous keys, and synchronous TLS APIs expose authentication work one connection at a time. We present WeaveTLS, a wire-transparent architecture that executes ML-DSA authentication across concurrent TLS connections....
Ethereum Proof-of-Stake (PoS) clients must verify large volumes of Boneh--Lynn--Shacham (BLS) signatures for attestations, sync-committee messages, and other consensus-critical objects within fixed slot deadlines. This recurring cost competes with state transition, fork choice, and message propagation for client CPU time, so reducing it increases the verification headroom available under bursty load. Prior cryptographic-engineering work has shown that SIMD can substantially accelerate...
In recent years, quantum circuit optimization has become an important research topic. Motivated by the fact that quantum gates act on fixed physical wires and modify only their target wires, we propose two SMT encodings: an exact-G encoding and an at-most-G encoding with null gates. Our method speeds up most tested 4-bit S-box instances, achieving up to approximately 130x speedup on the ELEPHANT S-box. Importantly, our method enables automated synthesis of practical 5-bit S-box quantum...
Post-quantum deployments need signatures that are both fast and small. ML-DSA gives a practical Fiat-Shamir lattice-signature baseline, but its signatures remain large enough to make bandwidth, certificate size, and signed-log storage first-order costs. Gaertner's iterative rejection sampling construction (CRYPTO'25) shows that this design family can be made much more compact. The open question is whether this theoretical design can be turned into a concrete, implementation-oriented...
Univariate-polynomial interactive oracle proofs (IOPs) over binary extension fields $\mathbb{F}_{2^m}$ underpin a class of plausibly post-quantum zkSNARKs, but rely heavily on polynomial arithmetic, where large-domain evaluation and division by subspace vanishing polynomials are the dominant prover costs. General-basis additive FFTs, such as Gao--Mateer and Lin-Chung-Han (LCH), accelerate the evaluation but impose a basis-conversion stage costing $O(n (\log n)^2)$ field additions and $O(n...
Two novel symmetric multidimensional affine nested variations of the Hill Cipher are presented. The Hill Cipher is a block polygraphic substitution encryption scheme based on a linear transformation of plaintext characters into ciphertext characters. In the time since Hill first published his encryption scheme, variations, modifications, and improvements of theoretical and practical importance have been published every year indicating that the Hill Cipher is an active area of cryptography...
This note collects, in compressed form, some techniques for modular multiplication with word-size (“short-limb”), or at most a-handful-of-words sized moduli as they are used in implementations of lattice-based cryptography: Barrett reduction and multiplication (in signed and unsigned flavors, with exact error, range, and canonicality analyses), Montgomery reduction and multiplication (including the folded-constant form, the precise equivalence with Barrett multiplication, even moduli,...
Rijndael-256 (R256), the 256-bit block variant of the Rijndael family, is practically relevant in ongoing NIST draft discussions on wider-block standardization and in several NIST post-quantum signature candidates. Relative to AES, R256 combines a wider $4\times8$ state with non-standard ShiftRows offsets $(0,1,3,4)$, invalidating key assumptions behind many AES-oriented optimizations. We study how these mismatches manifest on three targets and develop three corresponding adaptation...
Modern isogeny-based cryptosystems spend much of their running time in finite-field, elliptic-curve, and higher-dimensional isogeny arithmetic. Exploiting SIMD parallelism in these computations is nevertheless nontrivial: central routines such as Montgomery ladders contain loop-carried dependencies, while point, pairing, and theta-coordinate formulas expose only irregular fine-grained parallelism. We show that substantial SIMD parallelism can be recovered by reorganizing the arithmetic...
MAYO is a signature scheme based on the Unbalanced Oil and Vinegar (UOV) construction and a third round candidate in the NIST standardization process for additional post-quantum signature schemes. We present a memory-optimized pure- C implementation of MAYO signature verification that reduces RAM consumption by 97–99% compared to the reference implementation provided by the PQM4 project [KPR+] at the cost of increasing runtime by 50–200% and while maintaining code size. This reduction...
We present a formally verified implementation of the ML-KEM Number-Theoretic Transform (NTT) based on Plantard arithmetic, produced via a code generator that targets ML-KEM, ML-DSA, and FN-DSA from a single parameter triple. The generator embeds a static bound analyzer that places modular reductions at code-generation time without runtime branching, eliminating per-scheme manual tuning while preserving constant-time guarantees. Each generation produces structurally identical implementations...
Masking is a widely adopted countermeasure to protect cryptographic implementations from side-channel attacks. Subsequent research has focused on designing masking schemes and formally proving their security, notably through the development of automated tools, within models abstracting the reality of a sidechannel analysis. These designs rely on an external source of randomness; however, there is currently no consensus on the choice of (pseudo-)random number generators for masking. To the...
The post-quantum signature scheme Falcon (FN-DSA), currently being standardized by NIST as FIPS 206 (Initial Public Draft submitted August 2025, final standard expected 2026-2027), relies on a discrete Gaussian sampler whose critical bottleneck is the function fpr_expm_p63, computing $\lfloor \exp(-x) \cdot 2^{63} \rfloor$ for $x \in [0, \ln 2)$. While the reference implementation already employs a degree-12 fixed-point polynomial (FACCT), no segmented approximation has been studied for this...
Fully Homomorphic Encryption (FHE) has emerged as one of the key technologies for privacy-preserving computation, enabling arbitrary computation directly on encrypted data. Vectorized FHE schemes, such as Brakerski/Fan--Vercauteren (BFV), Brakerski--Gentry--Vaikuntanathan (BGV), and Cheon--Kim--Kim--Song (CKKS), are typically used in applications dealing with large datasets, for example, confidential database queries and private ML inference. These FHE schemes are based on the computational...
Boolean polynomial multiplication is the primary computational bottleneck of the Hamming Quasi-Cyclic (HQC) key encapsulation mechanism. In this paper, we reframe the Frobenius Additive FFT (FAFFT) in ring-theoretic terms, via quotient-ring homomorphisms and the Chinese Remainder Theorem. This perspective shows that a complete decomposition into evaluation points is unnecessary for multiplication, and naturally yields the Strided FAFFT (SFAFFT), which operates over smaller finite fields with...
Classical selection criteria for cryptographic S-boxes—nonlinearity $\mathrm{NL}$, differential uniformity $\delta$, boomerang uniformity $\beta_{\mathrm{B}}$, algebraic degree $\deg$—are invariants of affine equivalence. That property is exactly what blinds them to a class of side-channel weaknesses. The correlation-power-analysis (CPA) template distinguisher is governed by the Hamming-weight functional, and Hamming weight is not affine-invariant; it does not descend to the...
Background: The migration to post-quantum cryptography confronts resource-constrained Internet of Things (IoT) devices with a material performance cost. CRYSTALS-Dilithium, standardised as the Module-Lattice-Based Digital Signature Algorithm (ML-DSA) in FIPS 204, fixes the Keccak-based SHAKE functions as its only symmetric primitives, and profiling on embedded platforms identifies hashing as the largest single contributor to the scheme’s software cost. This review synthesises the performance...
This work composes SHA-2 and SHA-3 into a unified hardware architecture, bringing them together as a single, efficient cryptographic ensemble. This need is driven in particular by Post-Quantum Cryptography (PQC), where different standardized schemes rely on either SHA-2 or SHA-3/SHAKE primitives. Rather than enforcing strict round-level unification, the proposed design applies selective sharing across the most area-critical components, including a shared 25x64-bit register bank, shared...
We develop (mostly) the radix-2 number-theoretic transform (NTT) and its butterflies, the twisting trick and why it never changes the transform, the freedom to use Cooley--Tukey butterflies in both directions, incomplete NTTs, Good's trick, and the ways all of these combine---closing with the coefficient-bound bookkeeping that motivates the whole toolkit. This note is intended to help implementers of postquantum cryptography, and is compressed from the author's lecture...
We present a fast implementation of Falcon (FN-DSA) signature verification with AVX-512. On a modern AMD Zen5 core, it completes a Falcon-512 verification in 3.6 microseconds, 2.6 times faster than an already optimized baseline, with comparable gains on Zen4, and consistent results across clang 21 and gcc 15. The speedup comes from rewriting the Number-Theoretic Transform (NTT) and from vectorising all other stages of the verification algorithm. The novelty is to use a 32-bit...
Falcon was selected by NIST in 2022 for standardization as a post-quantum digital signature scheme. Among all standardized signature schemes, Falcon achieves the smallest signature size. Its main drawback, however, is its reliance on floating-point arithmetic, which plays a critical role in the security analysis. This reliance poses significant challenges for practical implementations: some platforms lack floating-point units, floating-point division is not constant time on many processors,...
Password-based authentication remains widespread, and large-scale sets of leaked hashes enable practical offline brute-force attacks. Multi-target attacks, which check candidates against large sets of hashes simultaneously, are particularly effective. Understanding the capabilities of low-cost platforms for such attacks is important to assess real-world password security risks. Therefore, we present BF², a modular and scalable FPGA–CPU framework that accelerates multi-target password...
Outsourcing computations to cloud providers raises significant data privacy concerns, making Privacy-Preserving Computation via Fully Homomorphic Encryption (FHE) increasingly vital. However, adapting data sorting routines to the FHE domain introduces severe performance bottlenecks. This survey systematizes the state-of-the-art in FHE-based sorting algorithms. A novel complexity metric, FHE-Effort, is introduced to accurately evaluate homomorphic circuit efficiency. Eighteen algorithms are...
Efficient polynomial multiplication and matrix-vector operations are fundamental to computational algebra and modern cryptography. In lattice-based post-quantum cryptography (PQC), schemes utilizing Number Theoretic Transform (NTT)-unfriendly rings require highly optimized subquadratic multiplication algorithms. In this paper, we establish a rigorous mathematical framework for generalized $k$-way split polynomial multiplication and Toeplitz Matrix-Vector Product (TMVP) algorithms over...
Hamming Quasi-Cyclic (HQC) is a code-based key encapsulation mechanism selected by NIST for standardization, making its resistance to implementation attacks critically important. We present a side-channel attack that exploits load/store leakage in the manipulation of HQC's sparse secret vectors. Analysing Cortex-M4 assembly generated from the reference implementation, we identify a leakage surface in which the low and high 32-bit halves of each 64-bit word...
Hybrid homomorphic encryption (HHE) lets a constrained client send compact symmetric ciphertexts while a server transciphers them into homomorphic ciphertexts, making HE-friendly ciphers such as PASTA a practical choice. Efficient and side-channel-secure execution of PASTA on embedded devices, however, remains challenging, since existing hardware relies on dedicated cipher cores and provides no side-channel protection. We present a hardware/software co-design of PASTA on RISQrypt, an...
Recent algorithmic advancements in the Multi-Party Computation-in-the-Head (MPCitH) paradigm have resulted in more efficient post-quantum digital signature schemes. MQOM is a MPCitH-based digital signature scheme and candidate in the ongoing NIST Post-Quantum Cryptography (PQC) standardization effort, offering performance competitive with lattice- and multivariate-based schemes in software. In this work, we develop a dedicated hardware accelerator for MQOM and analyze the impact of recent...
We give in this short note a circuit implementing the matrix-vector product with the 32x32 binary matrix of the AES MixColumns using 88 XOR gates. Previously known circuits minimizing this metric have been published in the past years and achieved 94 XOR, 92 XOR, 91 XOR, and 89 XOR. As far as we can tell, a circuit with 88 XOR was previously unknown.
SMAC is a recently proposed by Wang et al.~stand-alone Message Authentication Code (MAC) constructed from repeated applications of the AES round function and featuring an aggregation mode, SMAC-1$\times n$, for scalable parallel processing. Although originally designed for high-throughput CPU implementations leveraging AES-NI instructions, its structural properties suggest strong compatibility with hardware parallelism. However, no systematic FPGA-oriented architectural study of SMAC has...
This short note discusses in detail a folklore but little-known hybrid hash function grounded on both the discrete logarithm and short integer solution problems. In particular, specific satisfactory parameters are provided to ensure the standard $128$-bit security level for the lattice problem with $256$-bit module, which may be useful in its own right. The hash function is a natural generalization of the classical Pedersen and Ajtai ones. Nevertheless, to the authors' knowledge, no one has...
This paper proposes a high-speed hardware accelerator for QR-UOV, a multivariate scheme, that executes all three operations: key generation, signature generation, and signature verification. QR-UOV utilizes a quotient polynomial ring structure to reduce the public-key size of the original UOV scheme; however, this introduces functional requirements distinct from other multivariate schemes, such as polynomial-matrix operations over $\mathbb{F}_{q^\ell}$, coefficient expansion for the...
In reaction to the emerging quantum threat, the National Institute of Standards and Technology (NIST) seeks post-quantum secure digital signature schemes. NIST's ongoing competition recently advanced to the third round, in which the Unbalanced Oil and Vinegar scheme (UOV) is a promising candidate due to UOV's conservative design, small signatures, and performant signing and verification. While these benefits make UOV attractive, the implementation aspects for compact hardware acceleration of...
In this paper, we present an optimized implementation of Hamming Quasi-Cyclic (HQC) on the ARM Cortex-M4. We optimize (i) the polynomial multiplication and (ii) the support expansion in fixed-weight sampling, and (iii) propose an optional caching strategy that reuses the public transforms and hash recomputed under a fixed key. For the polynomial multiplication, the fixed-constant multiplications in the Frobenius additive FFT (FAFFT) butterfly spend nearly half of their instructions on VMOV...
The cost of quantum cryptanalysis is dominated by the quantum circuit of the target cipher. Estimating the quantum attack cost of a cipher thus requires building that circuit and measuring its qubit count, Toffoli count, and Toffoli depth. This is manual work that needs expert knowledge and must be redone for each cipher and each cost target. Large language models handle ordinary programming well, but their use in constructing quantum circuits for ciphers is still limited. In this work, we...
With the widespread adoption of Field Programmable Gate Arrays (FPGAs) in security-critical industries such as defense and telecommunications, ensuring the confidentiality of sensitive data processed by these devices has become paramount. Side-Channel Analysis (SCA) poses a significant threat, necessitating the protection of cryptographic primitives through effective and efficient countermeasures. Within the framework of well-established formal adversary models, Boolean masking offers...
The rise of quantum computing threatens widely deployed public-key cryptosystems, driving the adoption of post-quantum cryptography (PQC) algorithms that rely heavily on modular arithmetic. Existing hardware accelerators of the PQC algorithms for resource-constrained Internet-of-Things (IoT) devices remain limited and lack integrated fault detection mechanisms. In this work, we present CMALU, a Compact, fault-tolerant Modular Arithmetic Logic Unit supporting six operations on a single...
LESS is a code-based signature scheme built on the linear equivalence problem and, in its v2.0 round-2 form, a candidate in the NIST call for additional post-quantum signatures. No microcontroller implementation of it has been reported: the official benchmarking effort for the additional signatures excluded it on memory grounds, and an x86-massif cross-check puts the reference's peak stack at up to $\approx 836$~KB---beyond the SRAM of even the largest mainstream Cortex-M4. This paper...
To initiate migration to post-quantum cryptography (PQC), it is necessary to identify whether deployed software uses quantum-vulnerable (QV) public-key cryptographic schemes such as RSA, ECDSA, and Diffie–Hellman (DH). However, many ELF executables are distributed without source code, making it necessary to directly screen executable binaries for QV candidates. A prior tool, Quantum-vulnerable Executable Detection (QED), provides high precision but incurs substantial analysis cost, whereas...
Following the standardization of major post-quantum cryptography (PQC) algorithms, C implementations of ML-KEM, ML-DSA, and SLH-DSA have been rapidly deployed. However, known-answer tests (KATs) verify only functional correctness and do not establish the absence of timing leakage caused by secret-dependent branches, memory accesses, or variable-latency instructions. This paper presents CT-KAT, an integrated screening platform for assessing constant-time risks in PQC C implementations. CT-KAT...
AIMer is a post-quantum digital signature scheme with a conservative design. Its security relies only on the symmetric-key one-way function AIM2 and an MPC-in-the-Head (MPCitH) zero-knowledge proof. AIMer is a Korean post-quantum cryptography (KpqC) standard. However, the AIMer standard code released in January 2026 is a portable C reference implementation. It does not include processor-specific optimizations. As a result, it does not exploit AVX-512, a 512-bit vector instruction set...
As TLS 1.3 migrates to post-quantum cryptography (PQC), hybrid X.509 transition strategies—alternative-signature (Catalyst), Composite, Chameleon, and signature combiners—are compared on cost but rarely on whether they actually enforce the classical↔PQC binding they promise. We show they often do not, and that the failure persists even in stacks that do check the binding. The same BouncyCastle library accepts a Catalyst certificate carrying a forged ML-DSA signature on its default path yet...
HCTR2 is a wide-block encryption mode that encrypts one fixed-size message as a single unit, so that flipping a single plaintext bit re-randomizes the whole ciphertext. Its main use is disk encryption, where the message is a disk sector. We instantiate it with ARIA, the Korean national block-cipher standard, and implement it on an NVIDIA RTX 4080 GPU. With many independent messages, assigning one thread per message keeps the device occupied. At low queue depth, however, most of the GPU sits...
Primitive-only PQC benchmarks are insufficient for attributing composed hybrid costs on Cortex-M4 because shared hash backends, randomized-signature behavior, and fixed classical/wrapper work affect measured performance. We implement a common bare-metal Cortex-M4 harness for representative KpqC/NIST families, measuring uniform Hash-CT hybrid KEM benchmark rows with X25519 and Bindel et al. hybrid-signature AND-combiner rows. The goal is composed-cost attribution under a uniform benchmark...
This paper proposes an optimized GPU implementation of the ARIA-GCM authenticated-encryption pipeline (CTR keystream, GHASH authentication, and their AEAD composition): ARIA-CTR uses a packed 32-bit S-box staged in shared memory, GHASH is optimized separately with a fixed-key 4-bit Shoup lookup table, the two stages are integrated as both a two-kernel and a fused single-kernel AEAD, and the same aria_gcm.cu source is tuned for Ampere and Pascal through compile-time parameters. For ARIA-CTR,...
We present a quantum resource estimation of the Rijndael variants \[ N_b = N_k \in \{4,5,6,7,8\}, \qquad N_r = N_b + 6, \] under the NIST MAXDEPTH quantum cost model. Extending the AES quantum encryption oracle~\cite{ref5} parametrically to arbitrary $N_b = N_k$, we generalize the in-place key schedule, including the single- and double-\texttt{SubWord} cases, the \texttt{ShiftRows} offsets, and the round constants. We implement and verify the resulting oracles using ProjectQ. The...
Information set decoding (ISD) is the standard generic decoding attack considered for code-based cryptography. A concrete quantum-resource estimate for Grover-accelerated ISD requires an oracle whose dominant component is Gauss–Jordan elimination. We improve the elimination circuit of Perriello et al. [25] and Jang et al. [15] by not updating the entries that no later pivot or the final weight predicate reads. The required result vector is recovered by a parallel back-substitution on the...
This paper presents a memory-efficient and high-speed implementation of NTRU+, one of the key encapsulation mechanisms (KEMs) selected by Korea’s post-quantum cryptography project (KpqC), on the ARM Cortex-M4. NTRU+ is small enough to run on its own on a Cortex-M4 class microcontroller, yet in real embedded environments, the peak stack occupied by polynomial buffers and the running time dominated by the NTT become key constraints. To address this, in the proposed technique, we reduce memory...
FAEST is a symmetric-key post-quantum digital signature scheme and a third-round candidate in the NIST Additional Digital Signatures standardization process. Its signing path concentrates cost in two operations: round-wise constraint generation, which proves in zero knowledge that the AES circuit is computed correctly, and finite-field multiplication, which computes the leaf nodes of a vector commitment. This paper accelerates both operations on a CUDA-enabled GPU, with AES round constraint...
The objective of this work is to investigate methods for improving the self-tuning mechanism of ring oscillator (RO) based True Random Number Generators (TRNGs). It also examines the challenges involved in achieving a reliable and stable design over long-term operation. Furthermore, this work analyzes potential approaches to address these challenges and validates their effectiveness using the NIST statistical test suite.
Zero-knowledge proof systems are increasingly relying on the Sumcheck protocol to avoid the FFT-heavy structure of earlier SNARK designs. Sumcheck is well suited for GPU acceleration; it consists of sequential rounds where each round performs regular, parallelizable operations over large multilinear evaluation tables. The focus is on how to organize this work across rounds: intuitively, the active polynomial state should remain close to the device that processes it, the CPU-GPU boundary...
Multiplicative complexity have shown to be an important metric for efficient implementations in various contexts such as side-channel secure implementation and transciphering. We introduce a generic framework based on conjugacy to reduce the multiplicative complexity of block ciphers. Our approach exploits the iterative structure of the block cipher to build alternative implementation based on conjugate round operations with overall smaller multiplicative complexity. We apply this...
FrodoKEM, a key encapsulation mechanism based on the standard (unstructured) LWE assumption, is recommended as a conservative choice for post-quantum key exchange by agencies like BSI and ANSSI. As such, it has garnered substantial attention from an implementation security standpoint. In particular, several papers have looked into masking FrodoKEM, and, like for various other lattice-based cryptosystems, identified the Gaussian sampling operation as a major bottleneck. In FrodoKEM, it is...
Blockchain networks need to exchange messages and assets across independent systems, but cross-chain verification can expose private validation data to relayers, bridge logic, validators, or destination-chain components. This paper presents a prototype-based zero-knowledge verification layer for privacy-preserving blockchain interoperability. The prototype uses Circom and SnarkJS to generate Groth16 proofs, verifies those proofs in Rust using arkworks BN254, and maps the result into a...
The main objective of this paper is to acceler ate the post-processing of Quantum Key Distribution (QKD) using an energy-efficient pipelined architecture implemented on a Field-Programmable Gate Array (FPGA). The proposed architecture aims to improve processing speed while efficiently utilizing hardware resources. In addition, this work compares the proposed approach with existing approaches to demonstrate its performance and resource efficiency.