GPUThor: Amplifying Rowhammer Attacks via Non-Uniform Patterns to Exploit ECC-Protected GPUs
AuthorsChris S. Lin, Joyce Qu, Aditya Rajeev, Gururaj Saileshwar
Resources
GPUThor shows that carefully timed, non-uniform memory access patterns can turn ECC-protected NVIDIA GPUs into practical targets for powerful Rowhammer attacks.
Key results
GPUThor generated up to 23,500 times more bit flips than previous GPU Rowhammer techniques.
The attack pattern reaches 6.6 ACTs per tREFI.
The RTX A5000 reached 377K bit flips per GB.
The campaigns produced 387 double-bit flips that ECC could detect but not correct.
What the paper found
GPUThor shows that Rowhammer attacks against NVIDIA GPUs are far more powerful than earlier results suggested. The attack reverse-engineers GPU memory-access coalescing and GDDR6 Target Row Refresh behavior, then distributes repeated aggressor accesses across different warps and cachelines while using non-uniform, multi-tREFI patterns. Its selected pattern spans six tREFIs and reaches 6.6 ACTs per tREFI, concentrating activations on victim-adjacent rows while using decoys to evade mitigation. Across NVIDIA RTX A4000, A4500, A5000, and A6000 GPUs, GPUThor produces 72K to 377K bit flips per GB, or as much as 23,500 times more than prior GPU attacks. It also undermines SECDED ECC: experiments found 387 double-bit flips that ECC could detect but not correct, plus two triple-bit flips that caused silent data corruption. With ECC enabled on an RTX A6000, the attack triggered denial-of-service conditions at an average rate of one DUE per hour, and the researchers demonstrated privilege escalation by exploiting ECC miscorrection and the roughly 10-millisecond interval before the GPU is terminated. The findings challenge NVIDIA’s recommendation that enabling ECC is sufficient protection and motivate stronger error correction, row-activation tracking, and memory-integrity defenses.
Original abstract
GDDR memory in GPUs is vulnerable to Rowhammer attacks, where rapid memory accesses induce bit flips in adjacent cells, enabling data tampering and privilege escalation. However, prior GPU Rowhammer attacks trigger only tens to hundreds of bit flips, orders of magnitude fewer than CPU attacks, severely limiting their practical impact. This gap stems from the reliance of existing GPU Rowhammer attacks on uniform hammering patterns that activate aggressor and decoy rows equally, which results in low hammering intensity for aggressor rows. We present GPUThor, a high-intensity Rowhammer attack on NVIDIA GPUs leveraging non-uniform hammering. GPUThor reverse engineers GPU memory-access coalescing behavior to enable non-uniform hammering patterns on GPUs, that activate aggressor rows more intensely than decoy rows. Additionally, by identifying refresh instances when in-DRAM mitigations are applied, it constructs longer attack patterns that escape mitigation across refresh intervals, further increasing hammering intensity. Together, these techniques yield 500X to 23,500X more bit flips than prior GPU Rowhammer attacks, across several NVIDIA GPUs (A4000, A4500, A5000, A6000), reaching bit flip rates close to state-of-the-art CPU Rowhammer attacks. GPUThor also enables the first Rowhammer exploits on ECC-protected GPUs, inducing uncorrectable double and triple bit flips, making denial-of-service and privilege-escalation attacks practical even on GPUs with ECC enabled.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.