GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-Design
AuthorsYitong Ding, Jiawei Huang, Renyang Guan, Yangjie Zhou, Zihan Liu, Yu Feng, Shixuan Sun, Mingyi Guo, Jingwen Leng, Jian Weng
Resources
GPU-Tile-Sim models LLM GPU performance through dependency-aware tile graphs, offering a faster and more adaptable tool for hardware-software co-design.
Key results
Accuracy across GEMM and attention kernels on NVIDIA A100 and H100
Average end-to-end inference modeling error on H100
Geometric-mean simulation-speed improvement on A100 GEMMs
Average execution-time reduction compared with DRAM-based communication
Average reduction in NoC residence time
What the paper found
GPU-Tile-Sim, from researchers at Shanghai Jiao Tong University, the National University of Singapore, and KAUST, proposes GTSim, a simulation framework for NVIDIA GPU hardware-software co-design around modern LLM kernels. Instead of simulating individual instructions or using coarse analytical mappings, GTSim represents execution as a warp-centric directed tile graph: nodes describe tile-level computation, memory movement, or synchronization, while data edges and order edges capture dependencies, buffer reuse, warp specialization, and asynchronous overlap. Its TileLang IR frontend automatically extracts these graphs after software-pipeline and warp-role transformations, and its backend schedules ready nodes against throughput-oriented compute, memory, cache, DRAM, TMA, and on-chip NoC models. On NVIDIA A100 and H100 systems, GTSim models GEMM and attention kernels with MAPE from 1.22% to 6.50%, outperforming TileFlow and LLMCompass; for end-to-end Llama-3-8B inference, average MAPE reaches 8.71%. It is also 3.98 times faster than Accel-Sim on A100 GEMMs while retaining comparable accuracy. Design studies show that NoC-enabled fusion reduces execution time by up to 1.94 times, with an average 1.65 times speedup over DRAM communication, and topology-aware mapping cuts NoC residence time by 1.47 times on average. The framework extends to NVIDIA Blackwell by adding TMEM and tcgen05 semantics, demonstrating a practical path for evaluating architectures and kernels such as FlashAttention-4 without rebuilding an instruction-level simulator.
Original abstract
Modern LLM (large language model) workloads increasingly rely on optimized GPU kernels through hardware-software co-design. These kernels achieve high-performance through fine-grained dependency scheduling and computation-memory overlap. As such, they incur new challenges on existing GPU performance models. Instruction-driven simulators are costly to adapt to evolving architectures, while analytical models are too coarse to capture kernels' characteristics. We propose GPU-Tile-Sim, a tile-centric GPU simulation framework for LLM hardware-software co-design. The key insight is that modern LLM kernel performance is governed less by individual instruction latency than by the dependency structure that controls execution order and overlap. Accordingly, GTSim represents kernel execution as a warp-level tile graph whose nodes capture tile-level operations and whose edges encode data and ordering constraints. Using this representation, we design an automatic tile-graph frontend and a graph-driven simulation backend. We evaluate GTSim on representative GEMM, attention, and end-to-end LLM inference workloads. On A100 and H100 across both conventional and highly optimized kernels, GTSim achieves high performance-modeling accuracy (MAPE, Mean Absolute Percentage Error, 1.22%--8.71%). We further extend GTSim to Blackwell with preliminary validation, and demonstrate its effectiveness in analyzing software and architectural design choices.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.