SLAC: Access-Driven CPU-to-GPU Side-channel Attacks via System-Level Cache on Apple Silicon
AuthorsTianhong Xu, Saion K. Roy, Ruyi Ding, Aidong Adam Ding, Yunsi Fei
Resources
Researchers show that an unprivileged CPU process can spy on GPU workloads running on Apple Silicon by tracking footprints in the shared cache.
Key results
Apple M1 SLC cache sets resolved by reverse engineering
Throughput increase over CPrime+CProbe
Covert-channel throughput in bits per second
Covert-channel throughput in bits per second
Maximum input-keyword accuracy on TinyLlama and GPT-2 Medium evaluations
Maximum generated-response token accuracy on TinyLlama and GPT-2 Medium evaluations
What the paper found
SLAC demonstrates the first fine-grained, access-driven CPU-to-GPU Prime+Probe side channel against Apple Silicon’s shared System-Level Cache, allowing an unprivileged CPU process to observe GPU memory accesses without GPU co-residency. The work reverse-engineers Apple M1’s 4,096-set SLC, whose 12-bit set index is formed from hashed physical-address bits, and introduces collision-profile clustering to construct eviction sets despite Apple’s CPU-exclusive and GPU non-inclusive cache behavior. CPrime+CProbe performs both priming and probing on the CPU, while GPrime+CProbe uses GPU parallelism for priming; the latter provides a 6.4× throughput increase, reaching 400 Kbps versus 62.5 Kbps. Differential tracing and cache-set aggregation then enable end-to-end attacks on GPU-accelerated PyTorch workloads using Metal: a graph neural network attack reconstructs over 90% of edges across five datasets, and an LLM attack recovers input keywords with up to 94.8% accuracy and generated responses with up to 88.9% accuracy from TinyLlama and GPT-2 Medium. The results show that Apple’s unified-memory cache hierarchy can expose sensitive graph relationships, prompts, and model outputs across CPU-GPU security domains.
Original abstract
Modern heterogeneous System-on-Chip designs integrate CPU cores and a GPU that share a last-level cache (LLC) or system-level cache (SLC). This sharing exposes a new cross-domain attack surface, and existing attacks on integrated platforms either exploit coarse-grained cache-occupancy contention or require the adversary to co-reside on the GPU with the victim to obtain accurate timing measurements. In this work, we target Apple Silicon heterogeneous SoCs and discover that GPU memory accesses leave set-level footprints in the shared SLC, observable to an unprivileged CPU process. This keen observation enables the first fine-grained, access-driven, Prime+Probe-style CPU-to-GPU cache side-channel attacks against GPU workloads. We first reverse-engineer the Apple M1 SLC set-indexing functions and the interactions between local private caches and the SLC. Building on these findings, we construct the CPrime+CProbe SLC side-channel technique, which monitors GPU victim activity from the CPU at cache-set granularity. We then introduce an accelerated variant, GPrime+CProbe, in which an adversary leverages the GPU for faster SLC priming, yielding a 6.4x increase in the covert-channel throughput. Lastly, we demonstrate two end-to-end privacy attacks using the new side-channels: a graph-edge reconstruction attack on Graph Neural Networks (GNNs) that achieves 90% edge accuracy across five datasets, and an LLM privacy attack that recovers input keywords with up to 94.8% accuracy and model responses with up to 88.9% accuracy across TinyLlama and GPT-2 Medium models. Our results reveal a new class of microarchitectural vulnerabilities in Apple Silicon and call for secure system cache designs for heterogeneous SoCs.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.