Hardware
Building Faster Cryptography with Carryless Multiplication in NVIDIA CUDA 13.3
CUDA 13.3 introduces **clmad**, a hardware‑accelerated carryless multiply‑accumulate instruction available on all NVIDIA Ampere and newer GPUs (SM 80+), filling a long‑standing gap in GPU‑native binar
CUDA 13.3 introduces clmad, a hardware‑accelerated carryless multiply‑accumulate instruction available on all NVIDIA Ampere and newer GPUs (SM 80+), filling a long‑standing gap in GPU‑native binary field arithmetic. With clmad, GHASH in AES‑GCM achieves up to 18.8× speedup—roughly 6.3 TB/s on an NVIDIA B200—and similar throughput on GeForce RTX 5090. Additionally, sum‑check protocols for zero‑knowledge proofs over GF(2^128) are accelerated 3–13×, enabling large‑scale cryptographic and coding workloads on existing Ampere‑or‑later GPUs.
Source: NVIDIA Developer | 2026-07-15