NTH

GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-Design

AuthorsYitong Ding, Jiawei Huang, Renyang Guan, Yangjie Zhou, Zihan Liu, Yu Feng, Shixuan Sun, Mingyi Guo, Jingwen Leng, Jian Weng

July 31, 2026 2 min read
Watch on YouTube
The one-line take

GPU-Tile-Sim models LLM GPU performance through dependency-aware tile graphs, offering a faster and more adaptable tool for hardware-software co-design.

Key results

1.22%–6.50%
Kernel modeling MAPE range

Accuracy across GEMM and attention kernels on NVIDIA A100 and H100

8.71%
Llama-3-8B inference MAPE

Average end-to-end inference modeling error on H100

3.98
Speedup over Accel-Sim

Geometric-mean simulation-speed improvement on A100 GEMMs

1.65
NoC fusion average speedup

Average execution-time reduction compared with DRAM-based communication

1.47
Topology-aware mapping improvement

Average reduction in NoC residence time

What the paper found

GPU-Tile-Sim, from researchers at Shanghai Jiao Tong University, the National University of Singapore, and KAUST, proposes GTSim, a simulation framework for NVIDIA GPU hardware-software co-design around modern LLM kernels. Instead of simulating individual instructions or using coarse analytical mappings, GTSim represents execution as a warp-centric directed tile graph: nodes describe tile-level computation, memory movement, or synchronization, while data edges and order edges capture dependencies, buffer reuse, warp specialization, and asynchronous overlap. Its TileLang IR frontend automatically extracts these graphs after software-pipeline and warp-role transformations, and its backend schedules ready nodes against throughput-oriented compute, memory, cache, DRAM, TMA, and on-chip NoC models. On NVIDIA A100 and H100 systems, GTSim models GEMM and attention kernels with MAPE from 1.22% to 6.50%, outperforming TileFlow and LLMCompass; for end-to-end Llama-3-8B inference, average MAPE reaches 8.71%. It is also 3.98 times faster than Accel-Sim on A100 GEMMs while retaining comparable accuracy. Design studies show that NoC-enabled fusion reduces execution time by up to 1.94 times, with an average 1.65 times speedup over DRAM communication, and topology-aware mapping cuts NoC residence time by 1.47 times on average. The framework extends to NVIDIA Blackwell by adding TMEM and tcgen05 semantics, demonstrating a practical path for evaluating architectures and kernels such as FlashAttention-4 without rebuilding an instruction-level simulator.

Original abstract

Modern LLM (large language model) workloads increasingly rely on optimized GPU kernels through hardware-software co-design. These kernels achieve high-performance through fine-grained dependency scheduling and computation-memory overlap. As such, they incur new challenges on existing GPU performance models. Instruction-driven simulators are costly to adapt to evolving architectures, while analytical models are too coarse to capture kernels' characteristics. We propose GPU-Tile-Sim, a tile-centric GPU simulation framework for LLM hardware-software co-design. The key insight is that modern LLM kernel performance is governed less by individual instruction latency than by the dependency structure that controls execution order and overlap. Accordingly, GTSim represents kernel execution as a warp-level tile graph whose nodes capture tile-level operations and whose edges encode data and ordering constraints. Using this representation, we design an automatic tile-graph frontend and a graph-driven simulation backend. We evaluate GTSim on representative GEMM, attention, and end-to-end LLM inference workloads. On A100 and H100 across both conventional and highly optimized kernels, GTSim achieves high performance-modeling accuracy (MAPE, Mean Absolute Percentage Error, 1.22%--8.71%). We further extend GTSim to Blackwell with preliminary validation, and demonstrate its effectiveness in analyzing software and architectural design choices.

Read the original paper

More in AI Hardware

Browse all 34 papers →