Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability

arXiv:2610.01650 · cs.CR · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability".

Elias: Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability proposes a privacy-preserving federated learning framework that integrates homomorphic encryption for training utility with differential privacy…

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So we're looking at this paper today, "Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability." It seems like the core idea here is that they're merging two different privacy tools—homomorphic encryption for the training part and differential privacy for checking or releasing the model.

Elias: That sounds like a significant combination, Nadia; it suggests they aren't just tacking DP onto FL, but using HE to handle the actual training utility while reserving DP for specific inspection points. This approach is interesting because it tackles different privacy concerns with different mechanisms.

Priya: From a measurement standpoint, I'm curious what the actual data shows here; they claim this combined framework improves both model utility and the estimated privacy over a baseline that only uses differential privacy for training. That's a big claim to back up with solid metrics.

Nadia: Exactly, Priya; the abstract makes it clear that this method addresses the limitations of relying on either technique in isolation, offering a way to keep model utility while allowing private monitoring during training and secure downstream release, which is really important for sensitive decentralized data.

Elias: And from a cryptographic angle, they're leaning on Fully Homomorphic Encryption, specifically mentioning CKKS as a popular scheme that supports approximate arithmetic over fixed-point and complex numbers because it's generally preferred for machine learning applications.

Priya: I wonder how that FHE process integrates with the noise injection from differential privacy during the model inspection steps, because that's where things get complex and where utility can really suffer.

Nadia: That’s a fair point, Priya; they describe a workflow where training happens on encrypted updates, and then only separate copies of the global updates are perturbed with noise for inspection or release.

Elias: And that separation is key because it means the training process itself remains on non-perturbed encrypted global updates, which helps preserve the utility during the training phase before any noise is added.

Priya: So, if we look at their results, they compare three scenarios: a DP-only baseline, one using HE and DP for inspection and release of final updates, and another where HE is used for training but DP is only applied to the very final model.

Nadia: And what Priya mentioned earlier was that they found the third setting yields lower estimated privacy loss than the first while still providing better utility, which really supports the thesis of this paper.

Elias: The authors used a Markov chain Monte Carlo-based Bayesian estimation method for DP based on Membership Inference Attacks to estimate privacy, and they adapted that method by introducing a new attack definition and test statistics tailored for federated learning.

Priya: That sounds sophisticated; using MCMC sampling to estimate the full posterior distribution of the privacy parameters allows them to account for uncertainty in attack performance, which is something most studies don't do.

Paper summary: Nadia: I'm interested in the practical application of that estimation; how cheap would it be for an adversary to exploit this combined framework? That’s a question we need to keep coming back to, Elias.

Elias: Well, Nadia, the security assessment hinges on how robust those HE schemes are against inference attacks and what specific parameters in their FHE setup might leave room for exploitation.

Priya: Looking at the utility analysis, they show that in scenario A3—the HE-based training with DP applied only to the final model—at round five hundred it achieves a test loss of one point zero nine compared to two point three seven for the DP-only approach.

Nadia: That reduction in test loss is substantial, Priya; maintaining predictive accuracy while boosting privacy protection sounds like a very practical outcome for real-world deployment.

Elias: The authors also noted that the trend of better utility alongside stronger estimated privacy persists throughout the remainder of training in that setting.

Priya: What about the intermittent monitoring scenario, which is setting two? They found that even when inspecting every round, P=one scenario A2/three achieves a lower posterior mean of epsilon compared to A1 over the training rounds.

Nadia: So, this framework allows for monitoring without actually disturbing or modifying the encrypted training process itself, which is a huge practical win for decentralized systems.

Elias: That suggests that the separation of concerns between HE for computation and DP for inspection creates a pathway where one mechanism can serve multiple roles simultaneously in this context.

Priya: Overall, it seems like the main conclusion drawn from this paper is that combining HE with DP offers a better privacy-utility trade-off compared to using either technique alone, especially when you look at the specific results they presented in their comparison of scenarios.

Nadia: I think the title itself, "Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability," perfectly captures that core contribution, focusing on utility while enabling private inspection.

Elias: And when we look at the broader implications, this work suggests a more nuanced approach to privacy-preserving federated learning where one technique handles the heavy lifting of computation and another handles the release mechanism.

Priya: For me, it means that in applications dealing with highly sensitive decentralized data, we might not have to choose between having a perfectly accurate model or having strong privacy guarantees during deployment.

Nadia: That’s exactly the kind of practical result we want to hear, showing how these complex cryptographic tools can work together effectively in a setting where data is highly distributed.

Elias: The authors' choice of CKKS for its support of approximate arithmetic over fixed-point numbers shows they were specifically targeting the needs of machine learning tasks within the HE framework.

Priya: So, to wrap up what we've discussed about this paper, it really seems like this framework provides a more balanced way to handle the privacy challenges inherent in federated learning by strategically applying encryption and noise.

Conclusion: Nadia: So, we're wrapping up our discussion on "Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability." The authors really nailed combining those two techniques for model inspection and availability.

Elias: Indeed, Nadia, I find the title itself quite telling about what they’re trying to achieve by merging those specific cryptographic tools. It points toward a unified approach that handles both training utility and subsequent privacy releases.

Priya: And from my side as the measurement researcher, it suggests they've found a way to balance the need for model accuracy with maintaining strong privacy guarantees during the inspection phase of federated learning.

Nadia: Exactly, Priya; it’s about making sure that when we want to check what the model is doing without compromising its training data, we have a solid mechanism for that.

Elias: I think their authors are pushing a concept where you don't have to choose between keeping the computation perfectly secure and being able to release a usable model later on.

Priya: That’s really interesting because it moves beyond just protecting the raw training data; they’re thinking about the entire lifecycle of the model, from creation to deployment.

Nadia: It means that for decentralized systems handling sensitive information, there's a path forward where we can have both a functional model and auditable privacy checks happening concurrently.

Elias: And I wonder how robust this combined system is against an adversary who might try to probe the noise injection or the encryption scheme itself.

Priya: That’s definitely something we need to keep looking into, but it seems their framework provides a much stronger privacy estimate when you look at their results compared to using either method alone.

Nadia: So, in simple terms, this paper is proposing a way to train models securely while still allowing for private monitoring and safe model release later on.

Elias: It's a neat idea because it shows that different cryptographic tools can serve complementary roles instead of competing against each other in this setting.

Priya: And the impact could be significant for industries where data privacy is paramount, like healthcare or finance, where you need both accuracy and strict confidentiality.

Nadia: So, while the technical details are deep—with FHE and MCMC estimation—the core message is a practical path for building trustworthy models in a decentralized environment.

Ceren Yıldırım, Kamer Kaya, Sinan Yıldırım, Erkay Savas

Sabancı University

cs.CR

Submitted: 2026-10-01

Updated: 2026-10-01

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 80/100

The gist: Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability proposes a privacy-preserving federated learning framework that integrates

Key concepts

Fully Homomorphic Encryption (FHE)
A cryptographic technique that allows computations to be performed directly on encrypted data without decrypting it first. In this framework, it enables the server to aggregate encrypted local model updates while keeping the actual model parameters secret from everyone involved in the training process.
Differential Privacy (DP)
A mathematical guarantee that protects individual data privacy by adding carefully calibrated noise to a dataset or output. Here, DP is used specifically on separate copies of global updates to provide privacy during inspection and final model release.
Model Inspection and Availability
The goal of using this combined approach is to allow authorized parties to inspect the training process (monitoring) and later securely release the final model. This ensures that sensitive decentralized data remains protected throughout its lifecycle.
MCMC-based Bayesian Estimation
A statistical method used to estimate how private a system is by simulating many possible privacy outcomes. It helps researchers provide a robust, uncertain measure of the privacy loss (epsilon) rather than just a single, potentially misleading number.

Terminology

Summary

Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability proposes a privacy-preserving federated learning framework that integrates homomorphic encryption for training utility with differential privacy for model inspection and release. This approach is significant because it addresses the limitations of relying on either technique in isolation, offering a method to preserve model utility while enabling private monitoring during training and secure downstream release, which is crucial for real-world applications involving sensitive decentralized data.

The gist: Our results show that this method improves both model utility and estimated privacy over the baseline method that relies solely on differential privacy for training.

Framework Components and Workflow

The proposed protocol combines Fully Homomorphic Encryption (FHE) with Differential Privacy (DP). The training workflow operates on encrypted and non-perturbed model updates, leveraging HE to preserve model utility during federated training. Differential Privacy is then applied only to separate copies of the global updates that are subsequently made publicly available for inspection or downstream use.

The key steps in the protocol include:

  1. Each client generates a secret key share using CKKS and collectively generates a public key for multiparty homomorphic encryption (MHE).

  2. The server sends the encrypted global model update, denoted by [θ(r)], to selected clients.

  3. Each selected client homomorphically trains a local model on their dataset, resulting in an encrypted local model update [θ(r)c].

  4. The server homomorphically aggregates these encrypted local updates to produce the next global model update: [θ(r)] = X c∈C(r) P nc[θ(r)c] P nc'∈C(r).

  5. If model inspection is needed at round r, the server creates a separate copy of the encrypted global update [θ(r)], generates a Gaussian noise term V(r), and produces a noisy version: [˜θ(r)] = [θ(r)] + V(r).

  6. The clients cooperate to decrypt this noisy model [˜θ(r)], making it available for inspection. Crucially, the training process continues from the non-perturbed encrypted global update [θ(r)].

  7. For downstream release, a Gaussian noise term V(R) is injected into the final global model before decryption and public release: [˜θ(R)] = [θ(R)] + V(R).

Privacy Estimation Methodology

To estimate the privacy of this complex framework, the authors adopt a Markov chain Monte Carlo (MCMC)-based Bayesian estimation method for DP based on Membership Inference Attacks (MIAs), adapted to FL. This method is chosen because it estimates the full posterior distribution of privacy parameters, accounting for uncertainty in attack performance and avoiding overly confident privacy estimates.

The auditing procedure involves three main steps:

  1. Designing the attack: This involves introducing a new attack definition and corresponding test statistics tailored for the FL setting, inspired by methods like LiRA [18].

  2. Measuring attack performance and obtaining error counts: The attack must be repeated multiple times to estimate type I (false positive) and type II (false negative) error probabilities.

  3. Estimating privacy from the error counts: The MCMC-DP-Est algorithm uses MCMC sampling to estimate the joint posterior distribution of the privacy parameters, denoted as (ϵ, s).

Experimental Scenarios and Results

The study compares three FL scenarios: 1) a DP-only baseline scenario where noise is injected into local model updates before aggregation; 2) a scenario using HE and DP applied to separate copies of model updates for inspection and release; and 3) a scenario using HE for training with DP applied only to the final model before release.

The findings indicate that the third setting yields lower estimated privacy loss than the first while providing better utility. Furthermore, intermittent monitoring (the second setting) can be achieved without perturbing or modifying the encrypted training process, while still providing lower estimated privacy loss than that of the first setting. Specifically, for A2/3 (HE-based training with DP inspection/release), even when inspecting every round (P=1), it consistently achieves a lower posterior mean of ϵ compared to A1 over the training rounds.

Utility and Noise Analysis

The paper demonstrates that using HE for training and applying perturbations only to the final model can significantly maintain predictive accuracy while offering stronger estimated privacy guarantees compared to the DP-only alternative. In Figure 2(b), at round 500, A3 (HE + final DP release) achieves a test loss of 1.09, compared to 2.37 for the DP-only approach, while providing a stronger estimated privacy protection with an estimated posterior mean of the privacy parameter ε of 4.32, compared to 7.26 for the DP-only approach. This trend persists throughout the remainder of training.

Improvements for AI systems

Here are the specific improvements to AI systems that can be derived from this research, along with what those improved systems can achieve:


) Improved AI System Capabilities:

  1. A privacy-preserving Federated Learning (FL) framework that enables collaborative model training without sharing raw sensitive data, while simultaneously allowing for periodic, privacy-preserving inspection of the training process and subsequent controlled model release.

  2. A mechanism to estimate the true privacy parameters of complex ML models during training using a novel Markov Chain Monte Carlo (MCMC)-based Bayesian estimation method tailored for FL settings, providing robust estimates that account for uncertainty in attack performance.

  3. A high-utility training pipeline that avoids the utility degradation associated with purely Differential Privacy (DP) by leveraging Fully Homomorphic Encryption (HE) during the training phase, while still enabling privacy-preserving inspection and downstream release via DP on the final model.

) Specific Improvements and Functionality:

  1. The proposed framework allows for a trade-off between training noise and privacy guarantees, enabling researchers to select specific noise variances based on utility requirements (e.g., achieving lower test loss for higher privacy).

  2. The system can perform intermittent model monitoring during training (using the A2/3 scenario) without perturbing or modifying the core encrypted training process, offering a way to detect client-side anomalies, low-quality data contributions, or potential model drift in real-time while maintaining strong privacy guarantees comparable to or better than a DP-only baseline.

  3. The final model can be released for downstream tasks (inference) in cleartext form after being subjected to calibrated Gaussian noise during the release phase (A3), ensuring that the released model remains private for subsequent use, even when it has been trained using HE throughout its lifecycle.

  4. The privacy estimation technique provides a more accurate assessment of privacy loss compared to relying solely on theoretical bounds or simpler empirical methods, allowing for more informed decisions regarding the required level of noise and inspection frequency.

) Specific Outcomes for AI Applications:

  1. Researchers can collaboratively train highly accurate deep learning models (like CNNs used in the FEMNIST dataset) across decentralized sources (e.g., hospitals, edge devices) without exposing individual patient data to any single entity or external observer during the training process.

  2. The system can be deployed in high-stakes environments where continuous monitoring is necessary to ensure model integrity and prevent data poisoning or client-side bias, without compromising the confidentiality of the underlying local training data.

  3. The resulting trained models can be reliably deployed for inference, and their privacy level can be quantitatively verified before release using the framework's built-in estimation tools.

  4. This system enables a privacy budget management strategy where model utility is maintained by applying DP noise only at specific checkpoints (inspection/release) rather than continuously during the computationally intensive training rounds.

Sources

Related papers