NTH

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives

AuthorsSiyuan Shen, Anton Korzh, John Bachan, Tiancheng Chen, Arnav Goel, Ludwig Schneider, Pouya Kousha, Zhenhao He, Sylvain Jeaugey, Kamil Iskra, Nishank Chandawala, Jeff R. Hammond, Torsten Hoefler

July 24, 2026 2 min read
Watch on YouTube
The one-line take

A new class of ultra-low-latency GPU collectives brings distributed LLM inference and HPC communication within 7% of the hardware speed limit.

Key results

2.37
Small-message AllReduce latency

µs on 4 GB200 GPUs, compared with 11.0 µs for the NCCL ring

8.7%
Llama-3.1-70B ITL reduction

Inter-token latency improvement from the optimized AllReduce

1.404
Measured speed-of-light bound

µs for AllReduce on two GB200 GPUs

7%
Best speed-of-light overhead

Overhead above the hardware lower bound for small messages

13%
Maximum vLLM ITL reduction

Improvement on 4 GPUs across evaluated LLMs

What the paper found

Researchers from ETH Zurich and NVIDIA present a latency-first redesign of GPU collectives for workloads where tiny communication delays sit directly on the critical path, especially long-context, decode-heavy inference for models such as Llama-3.1-70B and DeepSeek-V3. Building on NVIDIA NCCL’s device-side APIs, the work replaces costly global memory barriers with barrier-free protocols using LL synchronization, sentinel polling, bidirectional communication with double buffering, symmetric memory, NVLink multicast, and a new two-shot LL128 atomic AllReduce. On 4 GB200 GPUs, the proposed NCCL kernel reduces small-message AllReduce latency from 11.0 µs for the NCCL ring to 2.37 µs, producing an 8.7% reduction in inter-token latency for Llama-3.1-70B. The measured hardware speed-of-light lower bound is 1.404 µs, and the best kernels reach approximately 7% overhead above that bound. Integrated into vLLM across dense, mixture-of-experts, and hybrid models, the optimized collectives reduce inter-token latency by up to 13% on 4 GPUs and 11% on 8 GPUs, with comparable throughput gains. The study also improves NVIDIA cuSOLVERMp on the Alps supercomputer, showing that latency-centric collective design benefits both AI serving and traditional HPC. The released low-latency API lets developers construct custom NCCL collectives without rebuilding synchronization and buffer-management mechanisms from scratch.

Original abstract

GPU collective communication is typically optimized for bandwidth, yet many emerging workloads are increasingly limited by latency. Long-context decode-heavy large language model (LLM) inference is a prime example, where serving large models requires multiple GPUs, and many small collectives lie directly on the critical path of token generation. Therefore, even microsecond of overhead can impact performance and cost. In this work, we study how to approach the hardware Speed-of-Light (SoL) lower bound for GPU collectives within a scale-up network. We identify key principles for near-optimal designs, including barrier-free synchronization and efficient use of symmetric memory and multicast. Building on NCCL's device-side API, we develop low-latency interfaces for constructing custom collective kernels and use them to implement new symmetric collectives in NCCL. Microbenchmarks show substantial latency reductions for small and medium messages, reducing overhead to within 7% of the absolute SoL lower bound. When integrated into real applications, these kernels improve inter-token latency and throughput in LLM inference and accelerate cuSOLVERMp, demonstrating benefits for both AI inference and traditional HPC workloads.

Read the original paper

More in AI Hardware

Browse all 34 papers →