NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems
AuthorsConor James Green, William Won, Tuan Ta, Bradford M. Beckmann
Resources
NUNA speeds up multi-GPU inference by routing communication and placing computation with awareness of where traffic travels inside large GPU systems.
Key results
Maximum speedup from NUNA-aware placement alone.
Maximum collective speedup when placement and routing are combined.
Geomean reduction in time per output token across evaluated LLMs and GPU counts.
Largest end-to-end inference improvement reported.
Geomean communication improvement for decode collectives.
What the paper found
NUNA, from AMD Research and Purdue University authors Conor James Green, William Won, Tuan Ta, and Bradford M. Beckmann, identifies a new bottleneck in multi-die GPU scale-up systems: non-uniform network access caused by physical distance among compute units, HBM stacks, and inter-GPU I/O ports. Profiling AMD Instinct MI210, MI355X, and MI300X systems found that remote-transfer latency can vary by almost 2×, while projected next-generation designs show up to 3× NoC-latency differences and 1.8× worst-case transfer slowdowns. The paper introduces NUNA-aware routing, or NAR, which hashes latency-sensitive traffic onto nearby I/O-port subsets, and NUNA-aware placement, or NAP, which maps communication threadblocks and memory pages to nearby compute units and HBM stacks. Evaluated with ASTRA-sim 3.0 across All-Gather, All-Reduce, and All-to-All collectives from 2 to 64 GPUs, NAP alone delivers up to 1.5× collective speedup, while NAP combined with NAR reaches 1.8×. End-to-end tests on 12 LLM architectures, including Meta’s Llama, OpenAI’s GPT-OSS, Mixtral, Qwen3, and DeepSeek models, reduce mean decode time per output token by 7%, with a maximum improvement of 28%; decode collectives improve by 1.56×. The results show that conventional NUMA locality is insufficient: latency-sensitive AI inference, especially token-by-token decode in systems such as vLLM and SGLang, requires jointly optimizing compute placement, memory placement, and network routing.
Original abstract
Graphics processing unit (GPU) architectures are growing in size to meet the increasing compute and memory requirements. As GPU sizes increase, intra-socket wire transfer delay increases significantly. While previous research has optimized for compute and memory locality within a socket, the spatial impact on inter-GPU communication has not been well-studied. We introduce the term non-uniform network access (NUNA) to describe this emerging optimization dimension in multi-GPU systems. We specifically focus on latency-sensitive collective communication, common in machine learning inference. First, we highlight the need for NUNA-aware routing (NAR), choosing optimized, spatially-aware inter-GPU paths in large scale-up network topologies. Second, we introduce NUNA-aware placement (NAP), placing threadblocks and data near I/O to optimize the inter-GPU traffic. We demonstrate that the NAP optimizations alone offer up to 1.5x collective speedups over a locality-unaware baseline. Combining NAP with NAR yields up to 1.8x faster collectives over the locality-unaware baseline. This leads to 7% mean (28% max) time per output token speedup in machine learning inference.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.