AI as a Compiler: Compiling Triton kernels without the Triton compiler
AuthorsFrançois Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
AffiliationsStanford University · EPFL
Resources
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Key results
Number of common kernels evaluated across Ada, Hopper, and Blackwell GPUs.
Number of kernels from recent machine-learning papers evaluated on B200.
Peak speedup over autotuned Triton from decoding packed binary weights directly into Tensor Core operands.
B200 speedup from assigning complete softmax rows to threads in tensor memory.
Size after extending Volta to support modern PTX features, up from 42 kLoC.
What the paper found
This paper introduces the Triton AI Compiler, or TAIC, an LLM-based agent that bypasses Triton’s conventional intermediate representations and translates fixed Triton kernels directly into NVIDIA PTX. Using GPT-6 Astra, Triton 3.6.0 baselines, and NVIDIA L40S, H100, and B200 GPUs, TAIC generates, assembles, numerically tests, profiles with Nsight Compute, and iteratively revises PTX while preserving the kernel ABI and launch contract. Across 12 common kernels, performance ranges from 0.83× to 3.34× that of autotuned Triton, with the largest gains coming from transformations unavailable to the standard lowering pipeline: decoding BitDelta’s packed binary weights directly into Tensor Core operands reaches 3.34×, assigning complete softmax rows to threads in Blackwell tensor memory reaches 1.37× on FlashAttention, and reusing overlapping convolution windows reaches up to 2.23×. On 10 kernels drawn from recent machine-learning papers, TAIC also reaches 1.55× on FlashSinkhorn and 1.10× on a Mamba-2 primitive. The system uses architecture-specific instructions including Ada’s mma.sync, Hopper’s wgmma and TMA, and Blackwell’s tcgen05 and tensor-memory accumulators. To make verification practical, the authors extend the Volta symbolic PTX verifier with asynchronous execution, packed formats, Tensor Core operations, barriers, and Blackwell mechanisms, increasing its trusted code base from 42 kLoC to 60 kLoC. The results suggest that AI can replace substantial compiler-backend engineering for specialized GPU kernels, but only if verification keeps pace with rapidly evolving NVIDIA architectures.
Original abstract
Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA PTX: an LLM agent translates Triton kernels directly into PTX. We build an environment that evaluates candidate PTX, and an agentic harness in which an LLM translates Triton kernels into PTX. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers, AI lowering achieves 0.83x-3.34x the performance of autotuned Triton. The largest gains come from transformations that Triton's lowering pipeline does not perform, such as decoding packed binary weights directly into Tensor Core operands (3.34x on BitDelta), assigning each thread a complete softmax row in tensor memory (1.37x on FlashAttention), and reusing overlapping convolution windows (up to 2.23x). These results rely on a robust evaluation harness with comprehensive verification support. We build on Volta, an existing PTX verifier, and substantially extend it to support modern GPU architectures by introducing support for Blackwell's tcgen05 Tensor Core interface. This requires modeling three architectural features: managed tensor memory, descriptor-based operand layouts, and asynchronous execution coordinated through commits, waits, memory barriers, and proxy fences. We discuss the challenges involved in formalizing them, as well as the current limitations. Our results suggest an emerging future in which AI compilers replace custom-written intermediate representations and checkers, reducing the time and engineering effort required to bring up software for new general-purpose and custom chips.
Read the original paperMore in AI Hardware
Browse all 34 papers →Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.
Coherent error threshold for quantum LDPC codes
Zhengyi Han, Yuanchen Zhao, Yijia Xu, Yixu Wang, Zi-Wen Liu
This work shows that quantum LDPC codes can still reliably correct coherent errors below a universal noise threshold, strengthening the foundations of scalable fault-tolerant quantum computers.