NTH

Why Do Prefetchers Fail? Let Agents Answer

AuthorsXiangfeng Sun, Ceyu Xu, Ningzhi Ai, Zeyu Zhu, Yiyang Yuan, Yuan Xie

August 22, 2026 2 min read
Watch on YouTube
The one-line take

Agents iteratively diagnose prefetching failures and synthesize specialized hardware components that reportedly outperform leading human-designed prefetchers.

Key results

61.1%
MoP IPC speedup

Geomean IPC speedup over no prefetching on held-out workloads

1.91B
Agent campaign tokens

DeepSeek V4 Pro tokens consumed by the autoresearch campaign

17
Prefetcher engine count

Three Alecto base engines plus fourteen agent-generated specialists

30
Training traces

SPEC CPU2006 and SPEC CPU2017 traces used during design

11
Held-out traces

Unseen memory-intensive SPEC CPU2017 traces used for final evaluation

110 KB
On-chip storage

Synthesized state budget for the complete MoP prefetcher

What the paper found

The paper presents a performance-anomaly-driven autoresearch system that lets AI agents improve a deployed hardware prefetcher by repeatedly asking why measured memory misses remain. In ChampSim, the flow localizes hot miss program counters, gives per-PC agents hardware logs, source code, and sliced traces, requires runnable minimal cases to validate diagnoses, clusters failures into pattern families, and uses evolutionary code generation to synthesize specialized RTL-compatible sub-prefetchers. Residual gating, shared duplicate suppression, and accuracy throttling allow new engines to accumulate without disrupting existing coverage. Starting from Alecto’s three-engine ensemble, the system builds the 17-engine Mixture of Prefetchers, trained on 30 traces from SPEC CPU2006 and SPEC CPU2017 and evaluated on 11 unseen memory-intensive SPEC CPU2017 traces. Using 1.91B tokens from DeepSeek V4 Pro, MoP achieves a 61.1% geomean IPC speedup over no prefetching, exceeding Alecto, Berti, and Pythia by 14.5%, 21.6%, and 23.6%, respectively. RTL synthesis reports 110 KB of on-chip storage and 0.0347 mm2 in a 6nm process, showing that closed-loop agentic discovery can produce a practical prefetcher rather than merely a simulator-optimized design.

Original abstract

Hardware prefetchers are crucial to processor performance, yet their design remains labor-intensive and expert-driven. Architects inspect execution and memory-access traces, identify patterns, translate them into online hardware heuristics, and evaluate them in simulation, often with no guarantee of improvement. Human experts cannot systematically inspect billion-instruction traces across diverse real-world workloads. We present a performance-anomaly-driven autoresearch flow that repeatedly asks why a deployed prefetcher fails and uses the diagnoses to construct the Mixture of Prefetchers (MoP). Each iteration localizes high-impact unexplained misses to program counters, gives agents hardware logs, source code, and sliced traces, validates diagnoses through runnable minimal cases, and synthesizes specialized sub-prefetchers for recurring pattern families. Measured performance and remaining anomalies feed subsequent iterations, enabling simulator-in-the-loop discovery beyond model priors. The campaign consumes 1.91 billion DeepSeek V4 Pro tokens. On SPEC CPU2006 and SPEC CPU2017, MoP achieves a 61.1% geomean IPC speedup over no prefetching, outperforming the human-designed Alecto, Berti, and Pythia prefetchers by 14.5%, 21.6%, and 23.6%, respectively. RTL synthesis in a 6nm library reports 110 KB of on-chip storage and 0.0347 mm^2 area. To our knowledge, this is the first empirical demonstration that an agent-driven hardware-design process can produce an RTL-practical prefetcher that outperforms state-of-the-art human designs on unseen workloads.

Read the original paper

More in AI Hardware

Browse all 34 papers →