Purlin: Separating Orchestration from the Datapath of Collectives
AuthorsOsayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
AffiliationsStanford University · NVIDIA & Stanford University
Resources
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
Key results
Maximum latency speedup over baselines for seven collectives
Maximum bandwidth improvement over baselines
Average improvement in offline LLM-serving interactivity
Average improvement in online LLM-serving interactivity
Maximum online interactivity improvement under overload
Pixel-identical prompt outputs with matched seeds
What the paper found
Purlin proposes a three-layer architecture for GPU collectives that separates stable collective semantics from orchestration and hardware-specific data movement. Developers describe transformations between packed, scattered, or transposed layouts using copy or reduction operations; the SNAC protocol—Stage, Notify, and Consume—derives synchronization and dependencies at compile time, while the Atom layer supplies tunable copy and reduce primitives specialized for NVIDIA A100, H200, and B200 GPUs. SNAC uses separate push-based latency and pull-based throughput paths, and Atom’s shared-memory pipeline targets each GPU’s bandwidth-delay product without rewriting orchestration for new hardware. Across seven collectives, Purlin delivers up to 5.14-fold lower latency and 4.50-fold higher bandwidth than baselines including NCCL, NCCLX, and MSCCL++. In SGLang, Purlin improves offline LLM-serving interactivity by 1.13-fold on average and online interactivity by 1.26-fold on average, reaching 2.85-fold under overload. The evaluation covers Qwen3.5-122B-A10B, DeepSeek-V4-Flash, DeepSeek-V4-Pro, and Qwen-Image across three GPU generations. For diffusion generation, latency improves by up to 1.13-fold, while output quality remains unchanged, including pixel-identical results for 100 Qwen-Image-Bench prompts. The central result is that reusable orchestration can coexist with hardware-aware tuning, portability, and application-specific communication without coupling every collective implementation to a single GPU generation.
Original abstract
Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics, orchestration (where and when data moves), and the datapath (how data moves). This coupling makes it costly to adopt new hardware mechanisms and customize communication for applications. We present Purlin, a scale-up communication framework that separates these concerns. At the top of Purlin, we specify collectives as a naming of an input and output layout and a copy or reduction operation. In the middle, we introduce a shared orchestration protocol, Stage, Notify, And Consume (SNAC), which derives coordination from these specifications. Below SNAC sits a hardware-specific datapath we call Atom, which implements two key data movement primitives for collectives: copy and reduce. This separation lets us customize collectives and adopt new hardware mechanisms while reusing orchestration via SNAC. We evaluate Purlin on A100, H200, and B200 GPUs. Across seven collectives, Purlin achieves latency speedups of up to 5.14x and bandwidth improvements of up to 4.50x over baselines. Integrated into SGLang, Purlin improves offline LLM serving throughput and interactivity by 1.13x on average and up to 1.37x over baselines. For online LLM inference, Purlin improves interactivity by 1.26x on average and up to 2.85x, with the largest gain occurring under overload. For diffusion image generation, Purlin reduces end-to-end latency by up to 1.13x.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.
Coherent error threshold for quantum LDPC codes
Zhengyi Han, Yuanchen Zhao, Yijia Xu, Yixu Wang, Zi-Wen Liu
This work shows that quantum LDPC codes can still reliably correct coherent errors below a universal noise threshold, strengthening the foundations of scalable fault-tolerant quantum computers.