NTH

NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems

AuthorsConor James Green, William Won, Tuan Ta, Bradford M. Beckmann

August 5, 2026 2 min read
Watch on YouTube
The one-line take

NUNA speeds up multi-GPU inference by routing communication and placing computation with awareness of where traffic travels inside large GPU systems.

Key results

1.5
NAP collective speedup

Maximum speedup from NUNA-aware placement alone.

1.8
NAP plus NAR collective speedup

Maximum collective speedup when placement and routing are combined.

7%
Mean decode TPOT improvement

Geomean reduction in time per output token across evaluated LLMs and GPU counts.

28%
Maximum decode TPOT improvement

Largest end-to-end inference improvement reported.

1.56
Decode collective speedup

Geomean communication improvement for decode collectives.

What the paper found

NUNA, from AMD Research and Purdue University authors Conor James Green, William Won, Tuan Ta, and Bradford M. Beckmann, identifies a new bottleneck in multi-die GPU scale-up systems: non-uniform network access caused by physical distance among compute units, HBM stacks, and inter-GPU I/O ports. Profiling AMD Instinct MI210, MI355X, and MI300X systems found that remote-transfer latency can vary by almost 2×, while projected next-generation designs show up to 3× NoC-latency differences and 1.8× worst-case transfer slowdowns. The paper introduces NUNA-aware routing, or NAR, which hashes latency-sensitive traffic onto nearby I/O-port subsets, and NUNA-aware placement, or NAP, which maps communication threadblocks and memory pages to nearby compute units and HBM stacks. Evaluated with ASTRA-sim 3.0 across All-Gather, All-Reduce, and All-to-All collectives from 2 to 64 GPUs, NAP alone delivers up to 1.5× collective speedup, while NAP combined with NAR reaches 1.8×. End-to-end tests on 12 LLM architectures, including Meta’s Llama, OpenAI’s GPT-OSS, Mixtral, Qwen3, and DeepSeek models, reduce mean decode time per output token by 7%, with a maximum improvement of 28%; decode collectives improve by 1.56×. The results show that conventional NUMA locality is insufficient: latency-sensitive AI inference, especially token-by-token decode in systems such as vLLM and SGLang, requires jointly optimizing compute placement, memory placement, and network routing.

Original abstract

Graphics processing unit (GPU) architectures are growing in size to meet the increasing compute and memory requirements. As GPU sizes increase, intra-socket wire transfer delay increases significantly. While previous research has optimized for compute and memory locality within a socket, the spatial impact on inter-GPU communication has not been well-studied. We introduce the term non-uniform network access (NUNA) to describe this emerging optimization dimension in multi-GPU systems. We specifically focus on latency-sensitive collective communication, common in machine learning inference. First, we highlight the need for NUNA-aware routing (NAR), choosing optimized, spatially-aware inter-GPU paths in large scale-up network topologies. Second, we introduce NUNA-aware placement (NAP), placing threadblocks and data near I/O to optimize the inter-GPU traffic. We demonstrate that the NAP optimizations alone offer up to 1.5x collective speedups over a locality-unaware baseline. Combining NAP with NAR yields up to 1.8x faster collectives over the locality-unaware baseline. This leads to 7% mean (28% max) time per output token speedup in machine learning inference.

Read the original paper

More in AI Hardware

Browse all 34 papers →