The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing
AuthorsStefan Scholze, Johannes Partzsch, Sebastian Höppner, Florian Kelber, Andreas Dixius, Marco Stolba, Sirine Arfa, Marc Berthel, Georg Ellguth, Jim Garside, Hector A. Gonzalez, Stephan Hartmann, Thomas Kiel-Hocker, Dongwei Hu, Matthias Jobst, Khaleelulla Khan Nazeer, Tim Langer, Chen Liu, Gengting Liu, Matthias Lohrmann, Mantas Mikaitis, Felix Neumärker, Amirhossein Rostami, Stefan Schiefer, Tilo Schubert, Delong Shang, Bernhard Vogginger, Yexin Yan, Steve Furber, Christian Mayr
Resources
SpiNNaker2 is a scalable neuromorphic chip that combines brain-inspired event-based computing with deep-learning acceleration for more flexible and energy-efficient AI.
Key results
ARM M4F-based processing elements integrated on one SpiNNaker2 chip
Maximum measured deep-network performance in high-performance mode
Maximum measured efficiency in high-efficiency mode
Test accuracy of the five-layer spiking neural network
Energy reduction versus high-performance operation while preserving 1-millisecond real-time processing
Current installation scale reported for the TU Dresden SpiNNcloud system
What the paper found
Researchers at Technische Universität Dresden and the University of Manchester, including SpiNNaker pioneer Steve Furber, introduce SpiNNaker2, a flexible neuromorphic chip designed to bridge spiking neural networks and conventional deep learning. Its 152 ARM M4F processing elements combine software programmability with 16×4 MAC-array machine-learning accelerators, exponential and logarithm units, local SRAM, dynamic voltage and frequency scaling, and an event router supporting payload-bearing multicast packets. For INT8 inference, SpiNNaker2 reaches 4.5 TOPS in high-performance mode and 2.7 TOPS/W in high-efficiency mode, while supporting more than 150000 neurons and 1.8B synaptic events/s at a 1-millisecond simulation step. On the DVS Gesture benchmark, a five-layer spiking model achieved 92.04% accuracy; automatic per-core power selection preserved real-time 1-millisecond operation while reducing energy by 28% versus the high-performance setting. The OctopuScheduler maps convolution, fully connected, and matrix-multiplication layers across processing elements, enabling comparisons with NVIDIA A100 and Jetson Orin Nano hardware, although off-chip DRAM traffic limits large DNN workloads. The platform also supports event-based learning: E-prop training on Google Speech Commands used 8× less energy than an NVIDIA V100 when device utilization was considered. SpiNNaker2 scales from robotic sensor systems to the TU Dresden SpiNNcloud installation, which currently contains 5 million cores, positioning the architecture as a programmable substrate for sparse, irregular, and hybrid brain-inspired computation.
Original abstract
In deep learning, efficiency gets more and more important to compensate for the ongoing growth in model sizes and applications. Neuromorphic hardware has long been advocated as an upcoming alternative to deep networks, taking inspiration from the brain for achieving unprecedented energy efficiency. However, demonstrations of these gains only recently began to grow in complexity and real-world applicability. With SpiNNaker2, we present a chip that bridges the gap between deep networks and neuromorphic computing and allows for flexible exploration of computing approaches that combine both worlds. It features 152 processing elements equipped with an ARM M4F processor and dedicated accelerators, an extended SpiNNaker routing fabric for scalable event-based communication and a range of external interfaces for system integration, including Gbit Ethernet and an LPDDR4 memory interface. We demonstrate performance and efficiency of the SpiNNaker2 chip for neuromorphic and deep network workloads, as well as novel event-based computing approaches. For deep network workloads, the chip achieves up to 4.5 TOPS in high performance mode and up to 2.7 TOPS/W efficiency in high efficiency mode for INT8 workloads. The chip supports spiking neural networks with >150000 neurons and >1.8 billion synaptic events/s when simulated with a 1 ms time step. Its low baseline power of less than 250 mW allows for efficiency even under varying workload conditions, allowing to explore sparse and event-based modes of computation. All this demonstrates the chip's capabilities as a universal hardware platform for scalable brain-inspired computing and its combinations with mainstream deep network approaches.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.