GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization
AuthorsZaid Khan, Justin Chih-Yao Chen, Jaemin Cho, Elias Stengel-Eskin, Mohit Bansal
Resources
This paper shows that language models can act like smart GPU performance predictors, helping search systems test more kernel ideas while spending less time on expensive hardware runs.
Key results
The surrogate evaluation set contains 480 Triton kernels across six GPU Mode tasks.
Reinforcement learning with a Brier-shaped objective improves speedup recovered over the untrained GPT-OSS-20B baseline.
The Brier-shaped reinforcement-learning objective lowers ECE relative to the untrained GPT-OSS-20B baseline.
What the paper found
GPU Forecasters introduces a new role for large language models from OpenAI, Google DeepMind, and DeepSeek-style frontier systems: not just kernel generators, but selective surrogates that predict GPU kernel speedups and know when to abstain. The authors from UNC Chapel Hill, AI2, Johns Hopkins, and UT Austin recast kernel evaluation as an eight-bin relative-speedup forecasting task over Triton and CUDA code, using NVIDIA A100 measurements as ground truth. On a held-out set of 480 Triton kernels across six GPU Mode workloads, off-the-shelf models already recovered 79.9% to 93.1% of available speedup under 1%–50% measurement budgets, with Gemini-3 Flash best on ranking but poorly calibrated and GPT-OSS-20B stronger after training. They then fine-tuned OpenAI’s GPT-OSS-20B with reinforcement learning using correctness plus calibration rewards, and found that a Brier-shaped objective improved both confidence reliability and search utility: speedup recovered rose by 2.06 percentage points over the untrained baseline, while calibration error fell by 0.106. In end-to-end PUCT search, the surrogate let the optimizer evaluate four times as many candidates per GPU budget and found faster kernels than the equal-budget baseline on four of six tasks. The paper’s broader claim is that calibrated LLMs can act as virtual GPU models, using uncertain forecasts to defer only hard cases to expensive hardware measurement.
Original abstract
GPU kernels are the workhorse of modern deep learning, and optimizing them (via evolutionary search or coding agents) usually requires repeated measurement on target hardware. While these measurements provide the ground-truth signal necessary for kernel search, they are costly, because each evaluation of a kernel requires compilation and repeated execution on a GPU. As improvements in LLM inference reduce the cost of writing novel kernels and LLM-driven searches scale to large search budgets, on-device evaluation becomes a bottleneck. To address this, we study how LLMs can serve as selective GPU surrogates for kernel evaluation, by forecasting the performance of proposed kernels. A useful surrogate should be accurate, and it should be selective, by knowing when it could be wrong, and deferring to the GPU. To evaluate surrogates, we measure whether their forecasts are accurate, calibrated, and practically useful for recovering fast kernels under limited GPU-measurement budgets. Next, we study whether reinforcement learning can improve forecast accuracy and confidence calibration. Our experiments demonstrate that LLMs can accurately forecast relative kernel performance, that their utility can be improved through reinforcement learning. Used inside a kernel search, the surrogate lets the search consider several times as many candidates under the same GPU evaluation budget, and that leads to finding faster kernels than an equal-budget baseline. These results suggest that LLMs can play a broader role in kernel optimization, by acting as virtual models of a GPU rather than solely as kernel generators for search.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.