RNG: Flat Datacenter Networks at Scale
AuthorsGiacomo Bernardi, Ratul Mahajan, C. Seshadhri, Enrico Carlesso, Chinchu Merine Joseph, Saurabh Kumar, Pavan Manikonda, Luiza Popa, Randy Ram, Steven Robinson, Elizabeth Tennent
Resources
This paper introduces a cheaper, scalable datacenter network design that uses quasi-random graph routing and novel cable-shuffling hardware to beat or match fat-tree performance in production.
Key results
RNG is reported to be up to 45% cheaper than equivalent fat trees with the same oversubscription ratio.
Spraypoint finds over 50 edge-disjoint paths for almost all endpoint pairs in a 64-degree fabric.
Spraypoint finds over 60 edge-disjoint paths for half of endpoint pairs in a 64-degree fabric.
For comparison, 8-shortest-path routing has a median of 5 edge-disjoint paths.
For comparison, 64-shortest-path routing has a median of 35 edge-disjoint paths.
What the paper found
“RNG: Flat Datacenter Networks at Scale,” from Amazon Web Services, presents the first production deployment of a flat expander-based datacenter fabric built from quasi-random graphs. The paper solves three long-standing blockers for expander topologies: scalable routing, practical cabling, and predictable design. Its Spraypoint protocol is a fully distributed, demand-oblivious routing scheme that uses ECMP plus controlled waypoint “spraying” to create close to the maximum number of edge-disjoint paths; in evaluation it finds over 50 edge-disjoint paths for almost all endpoint pairs and over 60 for half of them in a 64-degree fabric, versus a median of 5 for 8-shortest-path routing and 35 for 64-shortest-path routing. For physical deployment, RNG introduces ShuffleBox, a passive optical device that internally shuffles fiber connections so cabling complexity stays comparable to fat trees while preserving expander connectivity. The authors also derive analytic models for path length, oversubscription, and cost, showing that RNG can be 9–45 percent cheaper than equivalent fat trees and, at the same oversubscription ratio, often delivers higher throughput across clique, hub, and matching traffic patterns. In production benchmarks on Amazon’s server mesh and edge mesh, RNG matches fat-tree performance for multipath transport and storage workloads, despite variable path lengths. The key novelty is that Amazon is not just proposing an expander network; it has made one operational at scale and defaulted it for most workloads.
Original abstract
We design and deploy in production the first flat datacenter networks. Our design, called RNG, is based on quasi-random graphs. While the cost and fault-tolerance benefits of such topologies have been long known, their practical realization has been hampered by a lack of scalable routing and cabling approaches. RNG has a new distributed routing protocol that exploits the properties of random graphs to find a large number of edge disjoint paths between pairs of endpoints. It uses a novel passive optical device that internally shuffles cables, which makes its cabling complexity similar to that of fat trees. We show that RNG matches or exceeds the performance of fat trees for a range of traffic patterns, despite being up to 45% cheaper. RNG is now the default datacenter network for most workloads at Amazon.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.