Why Do Prefetchers Fail? Let Agents Answer
AuthorsXiangfeng Sun, Ceyu Xu, Ningzhi Ai, Zeyu Zhu, Yiyang Yuan, Yuan Xie
Resources
Agents iteratively diagnose prefetching failures and synthesize specialized hardware components that reportedly outperform leading human-designed prefetchers.
Key results
Geomean IPC speedup over no prefetching on held-out workloads
DeepSeek V4 Pro tokens consumed by the autoresearch campaign
Three Alecto base engines plus fourteen agent-generated specialists
SPEC CPU2006 and SPEC CPU2017 traces used during design
Unseen memory-intensive SPEC CPU2017 traces used for final evaluation
Synthesized state budget for the complete MoP prefetcher
What the paper found
The paper presents a performance-anomaly-driven autoresearch system that lets AI agents improve a deployed hardware prefetcher by repeatedly asking why measured memory misses remain. In ChampSim, the flow localizes hot miss program counters, gives per-PC agents hardware logs, source code, and sliced traces, requires runnable minimal cases to validate diagnoses, clusters failures into pattern families, and uses evolutionary code generation to synthesize specialized RTL-compatible sub-prefetchers. Residual gating, shared duplicate suppression, and accuracy throttling allow new engines to accumulate without disrupting existing coverage. Starting from Alecto’s three-engine ensemble, the system builds the 17-engine Mixture of Prefetchers, trained on 30 traces from SPEC CPU2006 and SPEC CPU2017 and evaluated on 11 unseen memory-intensive SPEC CPU2017 traces. Using 1.91B tokens from DeepSeek V4 Pro, MoP achieves a 61.1% geomean IPC speedup over no prefetching, exceeding Alecto, Berti, and Pythia by 14.5%, 21.6%, and 23.6%, respectively. RTL synthesis reports 110 KB of on-chip storage and 0.0347 mm2 in a 6nm process, showing that closed-loop agentic discovery can produce a practical prefetcher rather than merely a simulator-optimized design.
Original abstract
Hardware prefetchers are crucial to processor performance, yet their design remains labor-intensive and expert-driven. Architects inspect execution and memory-access traces, identify patterns, translate them into online hardware heuristics, and evaluate them in simulation, often with no guarantee of improvement. Human experts cannot systematically inspect billion-instruction traces across diverse real-world workloads. We present a performance-anomaly-driven autoresearch flow that repeatedly asks why a deployed prefetcher fails and uses the diagnoses to construct the Mixture of Prefetchers (MoP). Each iteration localizes high-impact unexplained misses to program counters, gives agents hardware logs, source code, and sliced traces, validates diagnoses through runnable minimal cases, and synthesizes specialized sub-prefetchers for recurring pattern families. Measured performance and remaining anomalies feed subsequent iterations, enabling simulator-in-the-loop discovery beyond model priors. The campaign consumes 1.91 billion DeepSeek V4 Pro tokens. On SPEC CPU2006 and SPEC CPU2017, MoP achieves a 61.1% geomean IPC speedup over no prefetching, outperforming the human-designed Alecto, Berti, and Pythia prefetchers by 14.5%, 21.6%, and 23.6%, respectively. RTL synthesis in a 6nm library reports 110 KB of on-chip storage and 0.0347 mm^2 area. To our knowledge, this is the first empirical demonstration that an agent-driven hardware-design process can produce an RTL-practical prefetcher that outperforms state-of-the-art human designs on unseen workloads.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.