Cocoon: A System Architecture for Differentially Private Training with Correlated Noises

arXiv:2510.07304 · cs.AR, cs.AI, cs.CR, cs.LG · Submitted 2025-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Cocoon: A System Architecture for Differentially Private Training with Correlated Noises".

Jane: Machine learning models pose significant privacy risks by memorizing training data, necessitating differential privacy (DP) techniques like DP-SGD, but these methods often degrade accuracy.

Tom: First, who's behind it and why it matters.

Paper summary: Jane: That's a crucial point because if the privacy technique adds noise, we have to pay for it in terms of speed and resources during training, and the paper clearly shows those overheads can be quite substantial.

Lu: They identify these specific problems as being especially problematic for two scenarios: models with embedding tables and also for massive scale models with billions of parameters.

Meng: That tells us they aren't just looking at a general problem; they’re focusing on where the practical slowdowns are most intense.

Lalam: And I wonder how this affects the real-world deployment of huge AI systems, because if these overheads are too high, it could make training models that need to be private simply unfeasible on current hardware.

Tom: Well, this paper introduces Cocoon as a hardware-software co-designed framework specifically built to tackle these overheads and speed things up.

Jane: It seems like they're not just tweaking the math; they are building an entire system around it with specific strategies to handle the complexity of correlated noises.

Lu: They propose several optimization strategies, like using Cocoon-Emb for large embedding tables by pre-computing noises before training starts, and then adding noise coalescing to manage storage size.

Meng: That sounds like a very smart way to handle the memory challenge by not storing everything at once.

Lalam: And they also mentioned hot/cold splitting for entries in those tables, which suggests they can be more efficient about where they spend their computational effort during training.

Tom: So, it's about making the pre-computation and noise addition as smart and compact as possible before the actual heavy lifting begins.

Conclusion: Jane: So, we've been talking about how Cocoon tackles the noise headaches in differentially private training for massive AI models, and now it's time to wrap up by looking at what this whole "Cocoon" thing really means for our world.

Lu: I think what's really impressive is how they managed to combine those complex correlated noise ideas with actual hardware and software design, which is something I’ve always dreamed about for making AI more robustly private.

Meng: From my side, it sounds like a practical solution to a real engineering headache; they took these theoretical challenges and built something that actually runs on real hardware without completely crippling performance.

Lalam: For me, the implication is huge because if we can train much larger models under strict privacy rules without it taking forever or destroying accuracy, then we can deploy incredibly powerful AI systems in ways that are truly trustworthy for society.

Tom: That’s a big picture view, Lalam; so when you strip away the technical jargon, Cocoon is essentially providing the blueprint for training enormous AI without needing to sacrifice either privacy or speed too badly.

Jane: Right, and the authors who wrote this paper really showed off their ability to bridge that gap between high-level theoretical research and actual system design with hardware acceleration.

Lu: They did a smart job of identifying those specific bottlenecks—like the overhead from storing noise history—and then using things like Cocoon-Emb and CocoonNMP to directly attack those issues at the architectural level.

Meng: That hardware component, CocoonNMP, seems like the real thing for handling the sheer scale of billion-parameter models that we’re dealing with right now in startups.

Lalam: It makes me think about how this could change how we develop AI across different industries; imagine secure medical AI or financial modeling that can be trained on massive datasets without compromising personal information.

Tom: So, it really boils down to making high-performance, private AI training a reality instead of just a theoretical hurdle.

Jane: And the next thing we'll look at is how these architectural choices might influence future research directions in differential privacy itself.

The Pennsylvania State University

cs.AR, cs.AI, cs.CR, cs.LG

Submitted: 2025-10-08

Updated: 2026-10-01

Importance score: 79/100

The gist: Machine learning models pose significant privacy risks by memorizing training data, necessitating differential privacy (DP) techniques like DP-SGD, but these methods often degrade accuracy.

Key concepts

Correlated Noises
Instead of adding independent random noise at every step, this method uses noises that are mixed across iterations. This allows later noise additions to partially cancel out earlier ones, which can improve the model's accuracy while still maintaining differential privacy guarantees.
Cocoon-Emb
This strategy optimizes training for models with large embedding tables by pre-computing all necessary correlated noises before training begins. It stores these pre-computed noises in a compact format, exploiting the sparsity of gradients to save space and time during the actual training process.
CocoonNMP
This is a custom near-memory processing device built into hardware, designed for large models. It allows past noise information to be stored and processed efficiently in secondary memory without needing frequent, slow transfers back to the main processor, significantly speeding up computations.
Noise Coalescing
To prevent storing too many pre-computed noises, this technique aggregates or coalesces the noise right before it is needed. Instead of adding separate noises to every entry in every iteration, only one equivalent, aggregated noise is added when an entry is accessed or training concludes.

Terminology

Summary

Machine learning models pose significant privacy risks by memorizing training data, necessitating differential privacy (DP) techniques like DP-SGD, but these methods often degrade accuracy. This paper introduces Cocoon, a hardware-software co-designed framework that utilizes carefully designed correlated noises to improve model accuracy while addressing the non-negligible memory and compute overheads associated with these mechanisms when training large models.

The Gist

Cocoon is a hardware-software co-designed framework for efficient training with correlated noises, accelerating models through pre-computing and storing correlated noises in a coalesced format (Cocoon-Emb) and supporting large models through a custom near-memory processing device (CocoonNMP).

Background on Correlated Noise Mechanisms

The paper addresses the limitation of standard DP-SGD, which adds independent Gaussian noise at each iteration. To improve accuracy, recent works employ correlated noises that are mixed across iterations, allowing later noises to partially cancel earlier ones. Mathematically, the correlated noise at iteration t is calculated as:

zˆt = (zt − min(t,ˆb−1)∑τ=1 C[t, t-τ]zˆt−τ)/C[t].

This calculation involves a weighted average of the past noise history using a mixing matrix C. The major additional overheads identified in the study are storing the noise history and performing GEMV operations between the stacked noise history and the mixing vector. These overheads become especially problematic for (1) models with large embedding tables and (2) large-scale, billion-parameter models.

Cocoon's Optimization Strategies

Cocoon is designed to mitigate these overheads through several key strategies:

  1. (Cocoon-Emb): This strategy accelerates training models with large embedding tables by pre-computing all the correlated noises for them before training, and storing the pre-computed noises in a compact format by exploiting their gradient sparsity. This involves:

  2. Noise Pre-computing: Instead of performing GEMV on each iteration, Cocoon-Emb pre-computes correlated noises for all the future iterations of the embedding tables before the actual training starts.

  3. Noise Coalescing: To solve the issue where storing all pre-computed noises is too large, Cocoon adds noise coalescing. This technique ensures that only an equivalent, aggregated or coalesced noise is added right before an entry is accessed or training ends, rather than adding separate noises to every entry in every iteration.

  4. Hot/Cold Splitting: Cocoon-Emb classifies each table entry as either hot or cold and only pre-computes and coalesces noise for cold entries, where the size of the coalesced noise is determined by the average number of entries that need noise to be added in each iteration.

Hardware Acceleration with CocoonNMP

For large models, Cocoon employs a custom near-memory processing (NMP) device called CocoonNMP. This hardware supports:

  1. Custom Near-Memory Processing (Cocoon-NMP): This device enables past noises to be stored and processed efficiently in secondary memory while avoiding frequent data transfers.

  2. CXL Memory Integration: The Cocoon-NMP prototype is implemented as an add-in card (AIC)-type custom board that integrates a CXL controller and an NMP engine into a Xilinx Versal FPGA. It features a custom GEMV engine that performs GEMV between a matrix stored in CXL memory and a vector provided by the CPU, amortizing vector transfer costs for reasonably large models.

System Characterization and Performance Results

The study characterized various ML models (CNN, ViT, LLM, DLRM) across different noise history sizes (up to 128) and band sizes (up to 64). The results demonstrated significant speedups:

  1. Cocoon-Emb improves the training time by "2.46–4.87× for ˆb > 8."

  2. CocoonNMP achieves a performance improvement of 1.55–3.06× over the baseline, depending on whether it is Cocoon-Emb or Cocoon-NMP, and the specific model/hardware setup (e.g., A5000 vs. A100 GPUs).

  3. For DLRMs, both GPU-GEMV and CPU-GEMV incurred nonnegligible slowdowns, up to 14.49× when the noise history was stored in main memory, which Cocoon addresses through its coalescing and hardware solutions.

  4. CocoonNMP consistently outperforms baselines by achieving a speedup of "1.55–2.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the Cocoon framework:

  1. Improve Differential Privacy (DP) Training Efficiency for Large Models:

  2. Reduce Training Latency and Increase Throughput for Large Language Models (LLMs) and Deep Learning Recommendation Models (DLRMs):

  3. Enable Efficient Fine-tuning of Massive Embedding Tables in Neural Networks:

  4. Enhance Privacy Guarantees Against Gradient Extraction Attacks While Maintaining High Model Accuracy:


Specific Improvements and Capabilities:

  1. Improve DP Training Efficiency for Large Models:

Inference with standard Differential Privacy (DP-SGD) is significantly slowed down by the overhead of managing correlated noise history, especially for large embedding tables or billion-parameter models, due to memory bottlenecks and high computational costs (GEMV operations). Cocoon introduces a hardware-software co-designed framework that mitigates these issues.

The improved system can train large models (e.g., LLMs) with correlated noise mechanisms at speeds of 1.55–10.82× over baselines, depending on the hardware used (FPGA/NMP). This allows for the training of models that are currently prohibitively expensive or slow to train under strict privacy constraints.

  1. Reduce Training Latency and Increase Throughput for Large Language Models (LLMs) and DLRMs:

The latency in training is often dominated by data transfer bottlenecks (between GPU, Main Memory, and secondary memory) and the computational cost of noise generation (GEMV). Cocoon-NMP addresses this by integrating a custom Near-Memory Processing (NMP) device onto the CXL memory controller. This allows the crucial correlated noise generation step to happen in parallel with GPU training while keeping data local to the NMP device, drastically reducing slow PCIe bus transfers.

The improved system can achieve 1.55–3.06× speedup when using Cocoon-NMP on a real prototype, resulting in faster iteration times and higher overall training throughput for models like GPT-2 or OPT, which are essential for rapid model development and fine-tuning of foundation models.

  1. Enable Efficient Fine-tuning of Massive Embedding Tables in Neural Networks:

Embedding tables (used heavily in DLRMs) create unique overheads during correlated noise training due to the need to store and process noise history related to every entry, leading to linear slowdowns with table size. Cocoon-Emb specifically targets this by implementing a pre-computing and coalescing strategy. Instead of calculating noisy gradients iteratively for every entry, it pre-computes aggregated noises for cold (infrequently accessed) entries and stores them compactly.

The improved system can fine-tune DLRMs with massive embedding tables with 2.33–10.82× speedup, allowing researchers to explore much larger parameter spaces and richer data representations without being bottlenecked by the noise history management overhead.

  1. Enhance Privacy Guarantees Against Gradient Extraction Attacks While Maintaining High Model Accuracy:

The core challenge is balancing strong privacy (DP) with accuracy, which is usually compromised by independent Gaussian noise in DP-SGD. Cocoon utilizes carefully designed correlated noises that allow earlier noises to partially cancel out later ones, maintaining better model accuracy than standard DP-SGD. Furthermore, the framework includes a noise coalescing technique that ensures only an aggregated noise is added when an entry is accessed, minimizing unnecessary additions and preserving the integrity of the privacy mechanism across all iterations.

The improved system provides a robust defense against attackers attempting to extract training data from final models and intermediate gradients (under Cocoon-Emb's weaker adversary model), offering a more practical path for deploying privacy-preserving ML services in sensitive domains like healthcare or personal data processing.

Related papers