NTH
Research collection

AI Hardware research

Research on hardware and systems for machine learning workloads. Follow findings on throughput, energy use, memory, and deployment constraints.

34 papers · Latest edition October 7, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All AI Hardware papers

Newest editions first.

04Hardware

Coherent error threshold for quantum LDPC codes

Zhengyi Han, Yuanchen Zhao, Yijia Xu, Yixu Wang, Zi-Wen Liu

This work shows that quantum LDPC codes can still reliably correct coherent errors below a universal noise threshold, strengthening the foundations of scalable fault-tolerant quantum computers.

Read analysis
09Hardware

Hardware-Aware FP4 FlashAttention-4

Robert Hu

This work redesigns FlashAttention for Blackwell’s FP4 hardware, achieving faster inference and training while showing that aggressively quantized distributed training can become unstable.

Read analysis
10Hardware

MaxKernel: Agentic Kernel Generation for TPUs

Shangkun Wang, Nina Cai, Charles Hoong, Julian Walker, Gerson Kroiz, George Vanica, Deepak Patil, Andi Gavrilescu, Hassan Sipra, Sethu Sankaran

MaxKernel uses collaborating AI agents and compiler feedback to automatically discover high-performance TPU kernels that can rival expert optimization.

Read analysis
12Hardware

Programmable cavity QED with a fiber-integrated atomic array

Stephan Roschinski, Johannes Schabbauer, Franz von Silva-Tarouca, Marvin Holten, Damien Bloch, Julian Léonard

Researchers combine individually controlled rubidium atoms with a fiber cavity to build a programmable platform for quantum networks and many-body quantum optics.

Read analysis
15Hardware

Exponential quantum advantage for learning signals with a single qubit

Ishaan Kannan, Sridhar Prabhu, Saeed A. Khan, Mandar M. Sohoni, Xingrui Song, Saswata Roy, Alen Senanian, Valla Fatemi, Peter L. McMahon, Jordan Cotler

A single controllable qubit dramatically cuts the measurements needed to learn classical signals, suggesting powerful near-term applications for quantum-enhanced sensing.

Read analysis
16Hardware

Why Do Prefetchers Fail? Let Agents Answer

Xiangfeng Sun, Ceyu Xu, Ningzhi Ai, Zeyu Zhu, Yiyang Yuan, Yuan Xie

Agents iteratively diagnose prefetching failures and synthesize specialized hardware components that reportedly outperform leading human-designed prefetchers.

Read analysis
25Hardware

The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing

Stefan Scholze, Johannes Partzsch, Sebastian Höppner, Florian Kelber, Andreas Dixius, Marco Stolba, Sirine Arfa, Marc Berthel, Georg Ellguth, Jim Garside, Hector A. Gonzalez, Stephan Hartmann, Thomas Kiel-Hocker, Dongwei Hu, Matthias Jobst, Khaleelulla Khan Nazeer, Tim Langer, Chen Liu, Gengting Liu, Matthias Lohrmann, Mantas Mikaitis, Felix Neumärker, Amirhossein Rostami, Stefan Schiefer, Tilo Schubert, Delong Shang, Bernhard Vogginger, Yexin Yan, Steve Furber, Christian Mayr

SpiNNaker2 is a scalable neuromorphic chip that combines brain-inspired event-based computing with deep-learning acceleration for more flexible and energy-efficient AI.

Read analysis
26Hardware

Don't Predict, Prioritize: Rethinking GPU Reliability Assessment

Difeng Ma, Changhua Pei, Yuanwei Lu, Quan Zhou, Zexin Wang, Yibo Zhu, Daxin Jiang, Dan Pei, Jingjing Li, Gaogang Xie

Instead of guessing when a GPU will fail, HeaRank identifies which GPUs are most likely to fail so operators can act before problems disrupt large-scale AI jobs.

Read analysis
27Hardware

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives

Siyuan Shen, Anton Korzh, John Bachan, Tiancheng Chen, Arnav Goel, Ludwig Schneider, Pouya Kousha, Zhenhao He, Sylvain Jeaugey, Kamil Iskra, Nishank Chandawala, Jeff R. Hammond, Torsten Hoefler

A new class of ultra-low-latency GPU collectives brings distributed LLM inference and HPC communication within 7% of the hardware speed limit.

Read analysis
34Hardware

RNG: Flat Datacenter Networks at Scale

Giacomo Bernardi, Ratul Mahajan, C. Seshadhri, Enrico Carlesso, Chinchu Merine Joseph, Saurabh Kumar, Pavan Manikonda, Luiza Popa, Randy Ram, Steven Robinson, Elizabeth Tennent

This paper introduces a cheaper, scalable datacenter network design that uses quasi-random graph routing and novel cable-shuffling hardware to beat or match fat-tree performance in production.

Read analysis