Shiftfly: Scaling the Accelerator Interconnect Past the Pod with a Shift-Routed Optical Tier
AuthorsEylon E. Krause
Resources
Shiftfly replaces a costly all-to-all optical tier with a compact shift-routed topology that scales accelerator clusters to hundreds of thousands of chips, while exposing tradeoffs in locality and slice allocation.
Key results
TPU 8i pod scale where Boardfly achieves a diameter of 7.
Matched optical budget used for both Boardfly and Shiftfly.
Chip count at which Shiftfly reduces worst-case distance from 23 to 19 hops.
Shiftfly’s reduction in worst-case chip hops at 400,000 chips.
Approximate Shiftfly spectral-gap advantage over Boardfly across tested scales.
Shiftfly’s remaining shared-content tree-cost advantage after locality-aware placement.
What the paper found
Shiftfly examines how to extend Google’s TPU interconnect beyond the 1,152-chip TPU 8i Boardfly pod. Boardfly uses complete graphs at every tier, delivering a chip-level diameter of 7, but its global tier would require 12,499 optical ports per group for a 400,000-chip machine, far beyond the 40-port budget. Shiftfly preserves Boardfly’s 4-chip building blocks and 32-chip groups, replacing only the global tier with a generalized Kautz digraph implemented as a fixed permutation on the existing optical circuit switch. Correct port accounting sets the Kautz out-degree to 20, yielding table-free, shift-register routing whose destination label defines the path and whose diameter scales logarithmically with group count. At 1,152 chips, Shiftfly loses—11 hops versus Boardfly’s 7—but at 400,000 chips it reduces worst-case distance from 23 to 19 hops, a 17% reduction, and delivers roughly 2.7× higher spectral expansion. The proposed in-network merging for shared inference content is far less valuable than expected: locality-aware placement accounts for 21.5% of Boardfly’s achievable tree-cost savings and 14.5% for Shiftfly, leaving only a 2.9% residual Shiftfly advantage. Operationally, failed-group replacement ties at 40 optical circuits, while global wiring requires 28 bits of Shiftfly control state versus 550,000 for Boardfly. The main regression is slice allocation: arbitrary Shiftfly subsets can be disconnected, so slices must be instantiated with optical-switch reconfiguration rather than carved out.
Original abstract
Google's TPU interconnect spent nine generations as a $k$-ary $n$-cube, whose diameter grows as $Θ(N^{1/n})$, before TPU 8i replaced it with Boardfly: a three-tier hierarchy in which every tier is a complete graph, giving pod diameter 7 over 1,152 chips. Boardfly suits a single inference pod but does not extend, because a complete global tier needs $G-1$ optical ports to reach $G$ groups. A 400,000-chip machine would need 12,499 per group against the 40 available, so another hierarchy level must be stacked, and each level costs four chip hops. We propose Shiftfly, which keeps Boardfly's building block and group verbatim and replaces only the global tier with a generalized Kautz digraph, installed as a fixed permutation on the optical circuit switch the fabric already owns. At an identical 40 ports per group, Shiftfly is flat, has guaranteed diameter $\lceil \log_d G \rceil$, routes without tables by a shift register, and addresses content natively. The evaluation is deliberately two-sided. Shiftfly loses at one-pod scale, where Boardfly achieves chip-level diameter 7 against Shiftfly's 11, and wins beyond it, cutting worst-case distance from 23 to 19 hops at 400,000 chips with roughly 2.7x better spectral expansion at equal cost. The redundancy argument that motivated the design does not survive its own evaluation: locality-aware placement supplies most of the achievable saving on shared content, leaving the shift algebra a 2.9% residual, and we report the metric inversion that conceals this. We also measure operability. Replacing a failed group costs the same 40 optical circuits in both designs; deriving the global wiring costs 28 bits of control-plane state against 550,000. Slice allocation is the one regression: an arbitrary induced subset of a shift graph is disconnected, so slices must be instantiated rather than carved.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.