Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives
AuthorsSiyuan Shen, Anton Korzh, John Bachan, Tiancheng Chen, Arnav Goel, Ludwig Schneider, Pouya Kousha, Zhenhao He, Sylvain Jeaugey, Kamil Iskra, Nishank Chandawala, Jeff R. Hammond, Torsten Hoefler
Resources
A new class of ultra-low-latency GPU collectives brings distributed LLM inference and HPC communication within 7% of the hardware speed limit.
Key results
µs on 4 GB200 GPUs, compared with 11.0 µs for the NCCL ring
Inter-token latency improvement from the optimized AllReduce
µs for AllReduce on two GB200 GPUs
Overhead above the hardware lower bound for small messages
Improvement on 4 GPUs across evaluated LLMs
What the paper found
Researchers from ETH Zurich and NVIDIA present a latency-first redesign of GPU collectives for workloads where tiny communication delays sit directly on the critical path, especially long-context, decode-heavy inference for models such as Llama-3.1-70B and DeepSeek-V3. Building on NVIDIA NCCL’s device-side APIs, the work replaces costly global memory barriers with barrier-free protocols using LL synchronization, sentinel polling, bidirectional communication with double buffering, symmetric memory, NVLink multicast, and a new two-shot LL128 atomic AllReduce. On 4 GB200 GPUs, the proposed NCCL kernel reduces small-message AllReduce latency from 11.0 µs for the NCCL ring to 2.37 µs, producing an 8.7% reduction in inter-token latency for Llama-3.1-70B. The measured hardware speed-of-light lower bound is 1.404 µs, and the best kernels reach approximately 7% overhead above that bound. Integrated into vLLM across dense, mixture-of-experts, and hybrid models, the optimized collectives reduce inter-token latency by up to 13% on 4 GPUs and 11% on 8 GPUs, with comparable throughput gains. The study also improves NVIDIA cuSOLVERMp on the Alps supercomputer, showing that latency-centric collective design benefits both AI serving and traditional HPC. The released low-latency API lets developers construct custom NCCL collectives without rebuilding synchronization and buffer-management mechanisms from scratch.
Original abstract
GPU collective communication is typically optimized for bandwidth, yet many emerging workloads are increasingly limited by latency. Long-context decode-heavy large language model (LLM) inference is a prime example, where serving large models requires multiple GPUs, and many small collectives lie directly on the critical path of token generation. Therefore, even microsecond of overhead can impact performance and cost. In this work, we study how to approach the hardware Speed-of-Light (SoL) lower bound for GPU collectives within a scale-up network. We identify key principles for near-optimal designs, including barrier-free synchronization and efficient use of symmetric memory and multicast. Building on NCCL's device-side API, we develop low-latency interfaces for constructing custom collective kernels and use them to implement new symmetric collectives in NCCL. Microbenchmarks show substantial latency reductions for small and medium messages, reducing overhead to within 7% of the absolute SoL lower bound. When integrated into real applications, these kernels improve inter-token latency and throughput in LLM inference and accelerate cuSOLVERMp, demonstrating benefits for both AI inference and traditional HPC workloads.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.