SLAC: Access-Driven CPU-to-GPU Side-channel Attacks via System-Level Cache on Apple Silicon

arXiv:2608.09075 · cs.CR, cs.AR · Submitted 2026-08-14 · Read on arXiv

Tianhong Xu, Saion K. Roy, Ruyi Ding, Aidong Adam Ding, Yunsi Fei

Northeastern University · Louisiana State University

cs.CR, cs.AR

Submitted: 2026-08-14

Updated: 2026-08-17

Comments: Accepted to ACM CCS 2026. 15 pages

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: SLAC: Access-Driven CPU-to-GPU Side-channel Attacks via System-Level Cache on Apple Silicon This paper presents SLAC, the first fine-grained, access-driven, Prime+Probe-style CPU-to-GPU cache

Terminology

Summary

SLAC: Access-Driven CPU-to-GPU Side-channel Attacks via System-Level Cache on Apple Silicon

This paper presents SLAC, the first fine-grained, access-driven, Prime+Probe-style CPU-to-GPU cache side-channel attack on Apple Silicon. The authors target Apple M-series heterogeneous System-on-Chip (SoC) designs, where CPU cores and a GPU share a last-level cache (LLC) or system-level cache (SLC). The paper states: "We target Apple Silicon heterogeneous SoCs and discover that GPU memory accesses leave set-level footprints in the shared SLC, observable to an unprivileged CPU process. This keen observation enables the first fine-grained, access-driven, Prime+Probe-style CPU-to-GPU cache side-channel attacks against GPU workloads."

The authors first reverse-engineer the Apple M1 SLC set-indexing functions and the interactions between local private caches and the SLC. They note: We first reverse-engineer the Apple M1 SLC set-indexing functions and the interactions between local private caches and the SLC. They overcome two challenges: the CPU-exclusive SLC residency (where the SLC is exclusive with respect to the CPU private caches, which means that CPU-side data accesses cannot reliably populate or locate SLC cache lines) and the fully hashed set indexing (where the SLC does not draw its set index from contiguous low-order physical address bits directly). They develop a collision-profile clustering method for eviction set construction, which replaces noise-sensitive iterative pruning with collision-profile clustering. The reverse engineering yields 4,096 SLC sets with a 12-bit set index, and they recover a functionally equivalent set of indexing functions using XOR-based hash functions over physical address bits 7–32. They also characterize the SLC replacement policy as LRU, stating eviction is consistently access-order-dependent: lines accessed earlier are replaced before those accessed more recently, which is consistent with an LRU replacement policy.

Building on these findings, the authors construct the CPrime+CProbe SLC side-channel technique, which monitors GPU victim activity from the CPU at cache-set granularity. CPrime involves a two-step priming procedure: "In the first step, the attacker issues one load to each of the 65,536 cache-line addresses in SLC ES, accessing them in sequence so that every block enters L2... In the second step, the attacker loads L2 ES, which contains 12 addresses for each affected L2 set, which fill the set entirely and push out the surviving SLC ES data blocks into the SLC. CProbe then re-accesses every memory block in SLC ES and times each access, producing an integer trace of length 4,096 (one entry per SLC set), with each element taking a value in 0, 1,..., 16 that represents the number of SLC misses observed for the corresponding set."

The authors also introduce an accelerated variant, GPrime+CProbe, in which an adversary leverages the GPU for faster SLC priming. GPrime uses two GPU kernels: "the first kernel writes to all SLC ES memory blocks to invalidate any stale CPU private-cache copies or SLC copies, and the second kernel reads the same SLC ES addresses to load these memory blocks into the SLC cache lines. This yields a 6.4× increase in covert-channel throughput. The paper reports: CPrime+CProbe achieves a throughput of 62.5 Kbps, while GPrime+CProbe reaches 400 Kbps."

The authors demonstrate two end-to-end privacy attacks using the new side-channels. The first is a graph-edge reconstruction attack on Graph Neural Networks (GNNs) that achieves 90% edge accuracy across five datasets. Specifically, the paper reports: "With CPrime+CProbe, the CPU-only attacker achieves strong results. Single-node recall is consistently high, typically above 94% and often exceeding 98%... The full-graph recovery stage substantially closes this gap, raising CPrime+CProbe precision above 86% on all datasets while keeping recall above 90%. GPrime+CProbe attacks achieve higher precision and recall because GPU-based priming yields a cleaner SLC state. Full-graph precision then exceeds 99% on Cora, Citeseer, and Pubmed."

The second is an LLM privacy attack that recovers input keywords with up to 94.8% accuracy and model responses with up to 88.9% accuracy across TinyLlama and GPT-2 Medium models. For input keyword recovery, the paper reports: Table 4 reports keyword recovery accuracy, 76–81% across the two datasets and two models for CPrime+CProbe, rising to 91–95% for GPrime+CProbe. For response recovery, the paper states: CPrime+CProbe attacks recover 70–85% of output tokens correctly across the two models and two datasets. GPrime+CProbe attacks improve the accuracy by 4–5%.

The paper concludes: "The broader takeaway is that the shared SLC in Apple's unified memory architecture is a potent cross-domain attack surface, enabling fine-grained leakage of sensitive GPU workloads to an unprivileged CPU process. As heterogeneous CPU-GPU SoCs increasingly host privacy-critical ML workloads, these findings motivate careful redesign of cache hierarchies and coherence protocols to mitigate access-driven side-channels across compute domains."

Improvements for AI systems

Improvements to AI Systems:

  1. Privacy-Aware Model Execution Scheduler
  • Improvement: Integrate SLAC’s findings into ML deployment frameworks (e.g., PyTorch, TensorFlow Serving) to dynamically schedule GPU kernel execution or insert dummy memory-access patterns that obfuscate SLC set-level footprints.

  • Capability: Prevents an unprivileged CPU process from reconstructing GNN graph structures or LLM input/output tokens via cache side-channels, even on Apple Silicon. The system can detect high-risk GPU workloads (e.g., graph inference, token generation) and automatically enable noise injection or cache-flush barriers.

  1. Cache-Index Randomization for Heterogeneous SoCs
  • Improvement: Modify OS/hypervisor-level memory management to randomize physical address-to-SLC-set mappings (e.g., per-process XOR keys) based on the reverse-engineered indexing functions.

  • Capability: Breaks the attacker’s ability to build stable eviction sets, reducing side-channel signal-to-noise ratio. The improved system can run GPU workloads with near-zero performance overhead while making CPU-to-GPU leakage statistically indistinguishable from random noise.

  1. Real-Time Side-Channel Attack Detector
  • Improvement: Train a lightweight anomaly detector (e.g., a small CNN or LSTM) on CPU-side cache-access timing traces (like those from CPrime+CProbe) to flag suspicious Prime+Probe patterns targeting SLC sets.

  • Capability: The AI system can continuously monitor unprivileged CPU processes for eviction-set construction or high-frequency SLC probing, triggering alerts or throttling the suspected process before sensitive GPU data (e.g., LLM prompts, GNN adjacency lists) leaks.

  1. Adaptive Cache Partitioning Policy
  • Improvement: Use reinforcement learning to dynamically partition SLC ways between CPU and GPU based on workload sensitivity and real-time attack risk, informed by SLAC’s LRU characterization.

  • Capability: The improved system can isolate GPU-critical data into non-shared cache partitions during inference, eliminating cross-domain leakage while maintaining performance for non-sensitive tasks. It learns to balance security and throughput from historical attack patterns.

  1. Secure ML Compiler Optimizations
  • Improvement: Extend ML compilers (e.g., TVM, XLA) to insert cache-flush or data-reshaping instructions that break the set-index alignment exploited by SLAC, based on the recovered XOR hash functions.

  • Capability: Automatically transforms any GPU kernel into a side-channel-resistant version without manual code changes. The compiled model runs with identical accuracy but with SLC access patterns that are uniformly distributed across all 4,096 sets, defeating collision-profile clustering.

  1. Cross-Domain Coherence Protocol Verifier
  • Improvement: Build a formal verification tool that uses SLAC’s reverse-engineered SLC behavior (exclusivity, hashed indexing, LRU) to test new cache coherence protocols for similar vulnerabilities before silicon tape-out.

  • Capability: AI-driven simulation can predict whether a proposed SoC design leaks GPU data to CPU processes, enabling hardware architects to fix flaws early. The tool can also suggest alternative indexing functions that are provably resistant to Prime+Probe attacks.

Abstract

Modern heterogeneous System-on-Chip designs integrate CPU cores and a GPU that share a last-level cache (LLC) or system-level cache (SLC). This sharing exposes a new cross-domain attack surface, and existing attacks on integrated platforms either exploit coarse-grained cache-occupancy contention or require the adversary to co-reside on the GPU with the victim to obtain accurate timing measurements. In this work, we target Apple Silicon heterogeneous SoCs and discover that GPU memory accesses leave set-level footprints in the shared SLC, observable to an unprivileged CPU process. This keen observation enables the first fine-grained, access-driven, Prime+Probe-style CPU-to-GPU cache side-channel attacks against GPU workloads. We first reverse-engineer the Apple M1 SLC set-indexing functions and the interactions between local private caches and the SLC. Building on these findings, we construct the CPrime+CProbe SLC side-channel technique, which monitors GPU victim activity from the CPU at cache-set granularity. We then introduce an accelerated variant, GPrime+CProbe, in which an adversary leverages the GPU for faster SLC priming, yielding a 6.4x increase in the covert-channel throughput. Lastly, we demonstrate two end-to-end privacy attacks using the new side-channels: a graph-edge reconstruction attack on Graph Neural Networks (GNNs) that achieves 90% edge accuracy across five datasets, and an LLM privacy attack that recovers input keywords with up to 94.8% accuracy and model responses with up to 88.9% accuracy across TinyLlama and GPT-2 Medium models. Our results reveal a new class of microarchitectural vulnerabilities in Apple Silicon and call for secure system cache designs for heterogeneous SoCs.

Sources

Related papers