AI-Assisted Design of a Post-Quantum Cryptographic Accelerator: A Deployed-Silicon Case Study
AuthorsJungmin Park, Eunha Kim, Wooseop Kim, Seongjoon Cho, Byungho Cha
Resources
An AI agent designed a post-quantum crypto chip, and rigorous randomized testing exposed and fixed a hardware verification gap that standard tests missed.
Key results
Randomized validation signings completed with zero escapes
Deployed-baseline adversarial soak checks completed with zero failures
Overall success rate across the agentic design campaign
Resource utilization of the shipped Kintex-7 XC7K160T build
Positive setup margin in picoseconds for the critical clock domain
What the paper found
This case study describes an agentic AI workflow that took a unified ML-KEM-768 and ML-DSA-65 post-quantum accelerator from Verilog RTL through PCIe bring-up and deployed operation on a Kintex-7 XC7K160T FPGA, including on-chip key custody, shared NTT/INTT and Keccak-f[1600] datapaths, and hardware-generated entropy. Anthropic’s Claude Code, using Claude-family models such as Opus 4.8, worked under human review with Vivado, Python FIPS golden references, and AMD/Xilinx XDMA PCIe infrastructure. The central security finding is that ML-DSA rejection sampling creates message-dependent control paths that fixed known-answer tests cannot fully exercise: a BRAM-latency bug accepted an invalid signature at reject-loop iteration 5 despite passing the complete KAT regression. The replacement validation gate combined byte-exact reference comparison with randomized adversarial soaking, covering 301,343 data-dependent signings within 779,945 checks with zero escapes; the shipped design was also byte-exact across all six FIPS operations. Across 232 logged experiments, the assistant achieved 71.6% task success, with performance declining from 77–85% for documentation and research to 50–53% for synthesis and hardware bring-up, indicating that observability of physical FPGA behavior—not simply task complexity—limited reliability. The final v89 build reached 98.5% slice occupancy and retained 49 ps of setup margin at 500 MHz. End-to-end throughput reached 1,857 ML-KEM encapsulations per second and 480 ML-DSA signatures per second, while OpenSSL integration demonstrated hardware-custodied TLS 1.3 CertificateVerify. The study’s claim is therefore not autonomous hardware authorship, but that artifact-centered validation can establish trust independently of whether RTL was produced by humans or AI.
Original abstract
Post-quantum migration is mandated on published timelines, and silicon that ships with a defect cannot be patched remotely. The standard acceptance gate cannot detect an entire class of ML-DSA defects. Signing resamples until a candidate meets its norm bounds, so the executed path varies with the message, whereas known-answer tests (KATs) sample fixed values and reach only the depths their seeds trigger. Our accelerator passed its full KAT regression while carrying a norm check that outran block-RAM latency, leaving each candidate's final coefficients unverified; the escape surfaced at reject-loop iteration 5. The blind spot lies in the instrument, not the engineer; care cannot remove it. We replace that gate. A byte-exact golden-reference oracle paired with randomized adversarial soak drives the rejection loop past any fixed vector, closing the gap: 301,343 data-dependent signings, zero escapes. Because the gate judges artifacts and never authors, trust becomes separable from authorship, making AI authorship an answerable question. We report 232 logged experiments in which an agentic large language model drove a unified ML-KEM-768 and ML-DSA-65 accelerator with on-chip key custody from RTL to PCIe bring-up on one Kintex-7 XC7K160T, shipped at 98.5% slice occupancy. Success was 71.6%, following a hardware-coupling gradient, 77-85% for documentation and research against 50-53% for synthesis and bring-up, which observability can explain: failure concentrates where corrective signals are physical-side only. That so unreliable an author produced an artifact byte-exact across all six FIPS operations -- its deployed baseline surviving the same 779,945-check zero-failure soak -- is the claim.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.