NTH

A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family

AuthorsRishi Shah, Rishav Shrestha

August 19, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that many AI-generated GPU kernels that pass standard tests are silently wrong, and proposes stricter contracts to catch those failures.

Key results

12
Verifier gates

Adversarial contracts test numerical behavior, precision, determinism, shapes, aliasing, and hardware resources.

2,638
Audited accepted kernels

Machine-generated kernels already accepted by the source system’s harness.

62.1%
Contract violation rate

Accepted kernels carrying at least one verifier-detected contract violation.

39.5%
Tolerance-free failure rate

Accepted kernels failing a check that no tolerance adjustment can excuse.

1,487
KernelBench rejected by verifier

Kernels accepted by KernelBench’s standard test but rejected by the contract-grade verifier.

0.0033
Native backward oracle error

Worst primary-config relative error for the Blackwell tcgen05 backward against an fp64 oracle.

What the paper found

This paper challenges the reliability of LLM-generated GPU-kernel benchmarks, especially systems evaluated with NVIDIA hardware and Triton. Its contract-grade verifier replaces a single fixed-shape allclose test with 12 adversarial gates covering NaN and infinity propagation, determinism, shape generalization, precision, reduction-order stability, aliasing, device placement, and real hardware limits. Auditing 2,638 kernels that the Dr. Kernel system had already accepted, it found 62.1% with at least one contract violation and 39.5% failing tolerance-free checks; KernelBench accepted 1,487 kernels that the verifier rejected, versus only 14 in the reverse direction. The result was supported by a 7/7 positive control, calibration tests, 98.5% agreement with KernelBench’s correctness code, and manual audits. The paper’s systems contribution is a native NVIDIA Blackwell tcgen05 backward for the gated-linear-recurrence family, covering GDN, KDA, GLA, SSD or Mamba-2, and linear attention, including the reverse-state scan that open implementations leave on a Triton fallback. It manages Blackwell’s 512-column Tensor Memory limit through allocation-once, fixed-offset accumulators, and a single release, and trains all five family members. Against an independent fp64 oracle, the assembled backward reaches 0.0033 worst relative error in its primary configurations and is bit-for-bit deterministic, although it remains slower than the flash-linear-attention baseline. The broader message is that correctness contracts—not loose tolerances—are essential for trustworthy GPU-kernel generation and optimization.

Original abstract

Systems that generate GPU kernels with language models report high correctness rates. Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it if the output is close to a reference. A kernel can pass that test and still be silently wrong. It can return an ordinary number where the true answer is a NaN or an infinity, differ from run to run, break when the shape changes, or accumulate in fp16 where the reference keeps an fp32 total. We build the instrument that checks correctness properly: a contract-grade verifier of twelve adversarial gates, each a property a correct kernel must satisfy, several of them tolerance-free, so no choice of threshold can explain a failure away. Aimed outward, the verifier audits 2,638 machine-generated kernels that a public system's own harness had already accepted as correct. It finds 39.5% broken beyond any tolerance argument and 62.1% carrying at least one violation. The field's standard test accepts 1,487 kernels the verifier rejects, against only 14 the other way. We defend the finding four independent ways: a 7/7 positive control, a threshold-calibration sweep, 98.5% agreement with the reference benchmark's own correctness code, and a stratified hand-audit. Aimed inward, the verifier judges a kernel of our own: the first native Blackwell tcgen05 training backward for the gated-linear-recurrence (GDN) family, including the reverse-state stage the field still runs on a fallback. We establish its correctness independently, against a double-precision oracle, and train five family members through it. The correctness signal behind reported progress in kernel generation is far weaker than the numbers suggest, and a set of tolerance-free contracts would close most of the gap.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis