NTH

AI as a Compiler: Compiling Triton kernels without the Triton compiler

AuthorsFrançois Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini

AffiliationsStanford University · EPFL

October 7, 2026 3 min read
Watch on YouTube
The one-line take

An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.

Key results

12
Common-kernel suite

Number of common kernels evaluated across Ada, Hopper, and Blackwell GPUs.

10
Recent ML kernels

Number of kernels from recent machine-learning papers evaluated on B200.

3.34
BitDelta speedup

Peak speedup over autotuned Triton from decoding packed binary weights directly into Tensor Core operands.

1.37
FlashAttention speedup

B200 speedup from assigning complete softmax rows to threads in tensor memory.

60 kLoC
Verifier trusted code base

Size after extending Volta to support modern PTX features, up from 42 kLoC.

What the paper found

This paper introduces the Triton AI Compiler, or TAIC, an LLM-based agent that bypasses Triton’s conventional intermediate representations and translates fixed Triton kernels directly into NVIDIA PTX. Using GPT-6 Astra, Triton 3.6.0 baselines, and NVIDIA L40S, H100, and B200 GPUs, TAIC generates, assembles, numerically tests, profiles with Nsight Compute, and iteratively revises PTX while preserving the kernel ABI and launch contract. Across 12 common kernels, performance ranges from 0.83× to 3.34× that of autotuned Triton, with the largest gains coming from transformations unavailable to the standard lowering pipeline: decoding BitDelta’s packed binary weights directly into Tensor Core operands reaches 3.34×, assigning complete softmax rows to threads in Blackwell tensor memory reaches 1.37× on FlashAttention, and reusing overlapping convolution windows reaches up to 2.23×. On 10 kernels drawn from recent machine-learning papers, TAIC also reaches 1.55× on FlashSinkhorn and 1.10× on a Mamba-2 primitive. The system uses architecture-specific instructions including Ada’s mma.sync, Hopper’s wgmma and TMA, and Blackwell’s tcgen05 and tensor-memory accumulators. To make verification practical, the authors extend the Volta symbolic PTX verifier with asynchronous execution, packed formats, Tensor Core operations, barriers, and Blackwell mechanisms, increasing its trusted code base from 42 kLoC to 60 kLoC. The results suggest that AI can replace substantial compiler-backend engineering for specialized GPU kernels, but only if verification keeps pace with rapidly evolving NVIDIA architectures.

Original abstract

Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA PTX: an LLM agent translates Triton kernels directly into PTX. We build an environment that evaluates candidate PTX, and an agentic harness in which an LLM translates Triton kernels into PTX. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers, AI lowering achieves 0.83x-3.34x the performance of autotuned Triton. The largest gains come from transformations that Triton's lowering pipeline does not perform, such as decoding packed binary weights directly into Tensor Core operands (3.34x on BitDelta), assigning each thread a complete softmax row in tensor memory (1.37x on FlashAttention), and reusing overlapping convolution windows (up to 2.23x). These results rely on a robust evaluation harness with comprehensive verification support. We build on Volta, an existing PTX verifier, and substantially extend it to support modern GPU architectures by introducing support for Blackwell's tcgen05 Tensor Core interface. This requires modeling three architectural features: managed tensor memory, descriptor-based operand layouts, and asynchronous execution coordinated through commits, waits, memory barriers, and proxy fences. We discuss the challenges involved in formalizing them, as well as the current limitations. Our results suggest an emerging future in which AI compilers replace custom-written intermediate representations and checkers, reducing the time and engineering effort required to bring up software for new general-purpose and custom chips.

Read the original paper

More in AI Hardware

Browse all 34 papers →
03Hardware

Coherent error threshold for quantum LDPC codes

Zhengyi Han, Yuanchen Zhao, Yijia Xu, Yixu Wang, Zi-Wen Liu

This work shows that quantum LDPC codes can still reliably correct coherent errors below a universal noise threshold, strengthening the foundations of scalable fault-tolerant quantum computers.

Read analysis