Java Cup
Inside Java

News and views from members of the Java team at Oracle

Faster Post-Quantum Cryptography with JDK Intrinsics

The JDK supports modern post-quantum cryptography: ML-KEM, specified in FIPS 203; ML-DSA, specified in FIPS 204; and the hash-based HSS/LMS algorithms described in RFC 8554. These algorithms are implemented in Java, which provides portability and a reliable functional baseline. For particularly performance-sensitive operations, however, HotSpot can replace selected Java methods with optimized, platform-specific machine code known as an intrinsic. This code can take advantage of CPU features such as SHA-3 acceleration, SHA-256 acceleration, vector instructions, and efficient modular-arithmetic operations.

An intrinsic does not permanently replace the Java implementation. The Java method remains the portable fallback. The @IntrinsicCandidate annotation identifies methods that HotSpot may replace when a suitable implementation is available on the current platform.

Not every method benefits from becoming an intrinsic. A good candidate is both computationally expensive and frequently used, and its platform-specific implementation must provide a meaningful improvement over the Java fallback.

ML-KEM: A Natural Fit for Intrinsics

ML-KEM has several methods annotated with @IntrinsicCandidate. They are in the hot paths for key generation, encapsulation, and decapsulation. The most promising operations have highly regular computation over polynomials with exactly 256 coefficients. This means the same small set of arithmetic steps can be applied to each coefficient:

Assembly code is not inherently faster than Java; these operations are good candidates because their structure lets modern CPUs exploit the same patterns directly and repeatedly.

ML-DSA Shares Much of the Same Work

ML-DSA and ML-KEM are both lattice-based cryptographic schemes and share several broad categories of polynomial operations:

The implementations are not interchangeable. ML-KEM uses the modulus 3329, while ML-DSA uses 8380417, and the schemes use different coefficient representations and reduction arithmetic. They perform related kinds of work, but HotSpot generally needs algorithm-specific intrinsic implementations. ML-DSA also has operations with no direct ML-KEM counterpart. For example, implDilithiumDecomposePoly, which separates polynomial coefficients into two components during ML-DSA processing, is an ML-DSA-specific non-NTT intrinsic candidate.

Measured Performance Improvements on JDK 28

The following throughput improvements compare benchmarks with applicable intrinsics enabled against the Java-only implementation. Intrinsic bars show the minimum-to-maximum relative throughput across the applicable algorithm parameter sets. For example, ML-KEM-512 requires fewer computational resources than ML-KEM-1024 and therefore achieves higher throughput.

ML-KEM relative throughput

A 218% speed-up means the intrinsic-enabled implementation is 3.18× as fast as the Java-only version—not 2.18× as fast.

ML-DSA relative throughput

HSS/LMS Benefits from SHA-256 Acceleration

HSS/LMS is a hash-based signature scheme. In the measured verification workloads, SHA-256 accounts for most of the execution time, making SHA-256 intrinsics especially valuable.

HMS/LMS verification relative throughput

These results show why carefully selected intrinsics matter: accelerating a small, frequently executed primitive can substantially improve the performance of an entire cryptographic algorithm while retaining Java as the portable fallback.

References

¹ Butterfly operation. A butterfly operation is the basic two-input, two-output step used in transforms such as the FFT and NTT. Given values a and b, it commonly produces: a′ = a + wb and b′ = a − wb where w is a precomputed transform constant. In an NTT, all arithmetic is performed modulo a prime.

Butterfly operation

² Linux-aarch64 test system:

³ Linux-x64 test system: