NTH

PassNet: Scaling Large Language Models for Graph Compiler Pass Generation

AuthorsYiqun Liu, Yingsheng Wu, Ruqi Yang, Enrong Zheng, Honglei Qiu, Sijun He, Tai Liang, Jingjing Wu, Yuhan Zhou, Yiwei Zhang, Dongyan Chen, Weihan Yi, Xinqi Li, Siqi Bao

June 7, 2026 2 min read
Watch on YouTube
The one-line take

PassNet reframes LLM-based compiler optimization as graph pass generation, backing it with a large dataset and benchmark that show LLMs can meaningfully speed up real-world workloads.

Key results

18086
PassNet-Dataset unique graphs

PassNet-Dataset contains 18,086 deduplicated computational graphs collected from 100K real-world PyTorch and PaddlePaddle models.

279K
Expanded training subgraphs

The dataset construction expands the graph corpus into about 279K hierarchical training subgraphs using recursive folding, prefix analysis, and shape/dtype generalization.

200
PassBench evaluation tasks

PassBench curates 200 long-tail tasks for evaluating compiler pass generation.

2060
Subgraph-level tests

Those 200 PassBench tasks comprise 2,060 subgraph-level tests in total.

43%
TorchInductor end-to-end slowdown rate

Profiling 9,526 real subgraphs shows that 43% experience end-to-end slowdowns under TorchInductor’s default pipeline.

37%
Best aggregate gap to TorchInductor

Across evaluated models, the best aggregate score on PassBench still trails TorchInductor by 37%.

What the paper found

PassNet, from Baidu, reframes LLM-assisted compiler optimization away from standalone GPU kernel synthesis and toward graph compiler pass generation, where the model must emit structured rewrites that plug directly into tensor compilers such as TorchInductor, TVM, XLA, and MLIR-based stacks. The paper profiles 9,526 real subgraphs and finds a long-tail optimization ceiling: 43% of workloads slow down end-to-end under TorchInductor’s default pipeline, even though individual long-tail cases can see up to 3× speedup. To attack this, PassNet-Dataset collects 18,086 deduplicated computational graphs from 100K PyTorch and PaddlePaddle models and expands them into about 279K training subgraphs using Recursive Folding for recurring motifs, Prefix Analysis for fusible intervals, and shape/dtype generalization. For evaluation, PassBench curates 200 long-tail tasks with 2,060 subgraph-level tests and introduces Error-aware Speedup Score, or ESt, which jointly measures correctness, compilation stability, and runtime speedup while resisting benchmark gaming via AST inspection, runtime dispatch interception, and reverse evaluation order. Across frontier and open-source models including GPT-5.4, Claude-Opus-4.6, Claude-Sonnet-4.6, GLM-5.1, MiniMax-M2.7, and Qwen3, the best aggregate score still trails TorchInductor by 37%, despite some passes achieving 3.02× speedup on MaskFormer and 2.90× on BGE-Reranker. Fine-tuning Qwen3-30B-A3B on just 3,899 PassNet trajectories lifts AS from 0.139 to 0.371, a 2.67× gain, showing the bottleneck is consistency and training data, not raw capability.

Original abstract

Modern tensor compilers such as TorchInductor deliver substantial speedups on mainstream models, yet face a systematic performance ceiling on long-tail workloads -- our profiling shows that 43% of real-world subgraphs experience end-to-end slowdowns under default compilation. While LLMs offer a path toward automated optimization, existing efforts focus on standalone kernel generation. We argue that pass generation -- where LLMs author structured graph transformations that integrate directly into compiler pipelines -- is the more appropriate abstraction. We propose PassNet, the first large-scale ecosystem for LLM-based compiler pass generation, comprising: (1) PassNet-Dataset, over 18K unique computational graphs from 100K real-world models; and (2) PassBench, 200 curated long-tail fusible tasks (comprising 2,060 subgraphs in total) evaluated under the Error-aware Speedup Score (ES_t) -- a metric unifying correctness, stability, and performance -- with layered integrity defenses against systematic LLM exploitation. Experiments reveal that PassBench is both highly discriminative and genuinely unsaturated: the best frontier model trails TorchInductor by 37% in aggregate, yet on individual subgraphs LLMs achieve up to 3x speedup over the same compiler -- indicating that the bottleneck is consistency, not capability. Fine-tuning a small model on merely ~4K PassNet trajectories yields a 2.67x improvement approaching frontier-model performance, demonstrating substantial headroom and validating PassNet as live training infrastructure for advancing LLM-driven compiler optimization. All data, benchmarks, and tooling are publicly available.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis