A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family
AuthorsRishi Shah, Rishav Shrestha
Resources
This paper shows that many AI-generated GPU kernels that pass standard tests are silently wrong, and proposes stricter contracts to catch those failures.
Key results
Adversarial contracts test numerical behavior, precision, determinism, shapes, aliasing, and hardware resources.
Machine-generated kernels already accepted by the source system’s harness.
Accepted kernels carrying at least one verifier-detected contract violation.
Accepted kernels failing a check that no tolerance adjustment can excuse.
Kernels accepted by KernelBench’s standard test but rejected by the contract-grade verifier.
Worst primary-config relative error for the Blackwell tcgen05 backward against an fp64 oracle.
What the paper found
This paper challenges the reliability of LLM-generated GPU-kernel benchmarks, especially systems evaluated with NVIDIA hardware and Triton. Its contract-grade verifier replaces a single fixed-shape allclose test with 12 adversarial gates covering NaN and infinity propagation, determinism, shape generalization, precision, reduction-order stability, aliasing, device placement, and real hardware limits. Auditing 2,638 kernels that the Dr. Kernel system had already accepted, it found 62.1% with at least one contract violation and 39.5% failing tolerance-free checks; KernelBench accepted 1,487 kernels that the verifier rejected, versus only 14 in the reverse direction. The result was supported by a 7/7 positive control, calibration tests, 98.5% agreement with KernelBench’s correctness code, and manual audits. The paper’s systems contribution is a native NVIDIA Blackwell tcgen05 backward for the gated-linear-recurrence family, covering GDN, KDA, GLA, SSD or Mamba-2, and linear attention, including the reverse-state scan that open implementations leave on a Triton fallback. It manages Blackwell’s 512-column Tensor Memory limit through allocation-once, fixed-offset accumulators, and a single release, and trains all five family members. Against an independent fp64 oracle, the assembled backward reaches 0.0033 worst relative error in its primary configurations and is bit-for-bit deterministic, although it remains slower than the flash-linear-attention baseline. The broader message is that correctness contracts—not loose tolerances—are essential for trustworthy GPU-kernel generation and optimization.
Original abstract
Systems that generate GPU kernels with language models report high correctness rates. Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it if the output is close to a reference. A kernel can pass that test and still be silently wrong. It can return an ordinary number where the true answer is a NaN or an infinity, differ from run to run, break when the shape changes, or accumulate in fp16 where the reference keeps an fp32 total. We build the instrument that checks correctness properly: a contract-grade verifier of twelve adversarial gates, each a property a correct kernel must satisfy, several of them tolerance-free, so no choice of threshold can explain a failure away. Aimed outward, the verifier audits 2,638 machine-generated kernels that a public system's own harness had already accepted as correct. It finds 39.5% broken beyond any tolerance argument and 62.1% carrying at least one violation. The field's standard test accepts 1,487 kernels the verifier rejects, against only 14 the other way. We defend the finding four independent ways: a 7/7 positive control, a threshold-calibration sweep, 98.5% agreement with the reference benchmark's own correctness code, and a stratified hand-audit. Aimed inward, the verifier judges a kernel of our own: the first native Blackwell tcgen05 training backward for the gated-linear-recurrence (GDN) family, including the reverse-state stage the field still runs on a fallback. We establish its correctness independently, against a double-precision oracle, and train five family members through it. The correctness signal behind reported progress in kernel generation is far weaker than the numbers suggest, and a set of tolerance-free contracts would close most of the gap.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.