SparLeak: Privacy Leakage from Sparse Attention in LLM Inference on Shared GPUs
cs.LG, cs.CR
Submitted: 2026-09-30
Updated: 2026-09-30
Terminology
Sources
- Spill The Beans: Exploiting CPU Cache Side-Channels to Leak Tokens from Large Language Models
- Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
- SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
- The Llama 3 Herd of Models
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
- Differentially Private Representation for NLP: Formal Guarantee and An Empirical Study on Privacy and Fairness
- Whisper Leak: a side-channel attack on Large Language Models
- GPUVM: GPU-driven Unified Virtual Memory
- Efficient Attention Mechanisms for Large Language Models: A Survey
- Gemma: Open Models Based on Gemini Research and Technology
- LLaMA: Open and Efficient Foundation Language Models
- Leaner Training, Lower Leakage: Revisiting Memorization in LLM Fine-Tuning with LoRA
- SpeContext: Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs
- XAttention: Block Sparse Attention with Antidiagonal Scoring
- On Protecting the Data Privacy of Large Language Models (LLMs): A Survey
- Qwen3 Technical Report
- Time Will Tell: Timing Side Channels via Output Token Count in Large Language Models
- NetEcho: From Real-World Streaming Side-Channels to Full LLM Conversation Recovery
- InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks