NTH

GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization

AuthorsZaid Khan, Justin Chih-Yao Chen, Jaemin Cho, Elias Stengel-Eskin, Mohit Bansal

June 5, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that language models can act like smart GPU performance predictors, helping search systems test more kernel ideas while spending less time on expensive hardware runs.

Key results

480
Held-out Triton kernels

The surrogate evaluation set contains 480 Triton kernels across six GPU Mode tasks.

2.06%
Brier training speedup recovered gain

Reinforcement learning with a Brier-shaped objective improves speedup recovered over the untrained GPT-OSS-20B baseline.

0.106
Brier training calibration error drop

The Brier-shaped reinforcement-learning objective lowers ECE relative to the untrained GPT-OSS-20B baseline.

What the paper found

GPU Forecasters introduces a new role for large language models from OpenAI, Google DeepMind, and DeepSeek-style frontier systems: not just kernel generators, but selective surrogates that predict GPU kernel speedups and know when to abstain. The authors from UNC Chapel Hill, AI2, Johns Hopkins, and UT Austin recast kernel evaluation as an eight-bin relative-speedup forecasting task over Triton and CUDA code, using NVIDIA A100 measurements as ground truth. On a held-out set of 480 Triton kernels across six GPU Mode workloads, off-the-shelf models already recovered 79.9% to 93.1% of available speedup under 1%–50% measurement budgets, with Gemini-3 Flash best on ranking but poorly calibrated and GPT-OSS-20B stronger after training. They then fine-tuned OpenAI’s GPT-OSS-20B with reinforcement learning using correctness plus calibration rewards, and found that a Brier-shaped objective improved both confidence reliability and search utility: speedup recovered rose by 2.06 percentage points over the untrained baseline, while calibration error fell by 0.106. In end-to-end PUCT search, the surrogate let the optimizer evaluate four times as many candidates per GPU budget and found faster kernels than the equal-budget baseline on four of six tasks. The paper’s broader claim is that calibrated LLMs can act as virtual GPU models, using uncertain forecasts to defer only hard cases to expensive hardware measurement.

Original abstract

GPU kernels are the workhorse of modern deep learning, and optimizing them (via evolutionary search or coding agents) usually requires repeated measurement on target hardware. While these measurements provide the ground-truth signal necessary for kernel search, they are costly, because each evaluation of a kernel requires compilation and repeated execution on a GPU. As improvements in LLM inference reduce the cost of writing novel kernels and LLM-driven searches scale to large search budgets, on-device evaluation becomes a bottleneck. To address this, we study how LLMs can serve as selective GPU surrogates for kernel evaluation, by forecasting the performance of proposed kernels. A useful surrogate should be accurate, and it should be selective, by knowing when it could be wrong, and deferring to the GPU. To evaluate surrogates, we measure whether their forecasts are accurate, calibrated, and practically useful for recovering fast kernels under limited GPU-measurement budgets. Next, we study whether reinforcement learning can improve forecast accuracy and confidence calibration. Our experiments demonstrate that LLMs can accurately forecast relative kernel performance, that their utility can be improved through reinforcement learning. Used inside a kernel search, the surrogate lets the search consider several times as many candidates under the same GPU evaluation budget, and that leads to finding faster kernels than an equal-budget baseline. These results suggest that LLMs can play a broader role in kernel optimization, by acting as virtual models of a GPU rather than solely as kernel generators for search.

Read the original paper

More in AI Hardware

Browse all 34 papers →