Privacy Leakage via Output Label Space and Differentially Private Continual Learning

arXiv:2411.04680 · cs.LG, cs.CR · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Privacy Leakage via Output Label Space and Differentially Private Continual Learning".

Jane: The paper was written by Authors not visible in provided pages. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

The core problem: Tom: So, the paper starts by setting up this scenario where a model is learning from two different datasets, D t and D't, but one extra sample makes it all different. That's the trigger for our discussion today.

Jane: It’s not just that the data is slightly different; it’s how those differences manifest in the classifier's output label space that is dangerous. The paper shows that an attacker can use prior knowledge about those small differences to guess exactly which training dataset was used, and they get one hundred percent accuracy at worst-case attack.

Meng: That level of accuracy is terrifying for any system designed to be robust; the security completely collapses because we aren't tracking this subtle leakage in the output labels.

Lu: The real challenge here is that this specific leak isn's just a sample problem, it’s a problem of "task-wise adjacency," meaning that in continuous learning, even if we only look at one task at a time, the way these tasks connect can still allow for this massive privacy violation.

Lalam: This points to how often we fail to account for the entire lifecycle of data, not just focusing on individual records but on the whole sequence of tasks. It's a cultural blind spot that needs to be addressed in our approach to AI design.

The mechanism and why CL is hard: Tom: Building on what we just discussed, the paper explains that this side-channel becomes especially relevant when data streams are continuous, which is what continual learning describes. The authors show how the output label space needs to change as new labels appear in a sequential stream of tasks.

Jane: Because these tasks are happening one after, D t is constantly evolving, and since we don't know all the possible classes upfront—the full label space Y—the standard way of just releasing the actual data labels is completely non-DP.

Meng: The authors even compare their method to existing approaches like Desai et al., who rely on memory buffers to mitigate forgetting. It’s interesting that this paper finds solutions without those buffers, which simplifies the practical deployment significantly.

Lu: The theoretical complexity comes from the "catastrophic forgetting" problem; we're not just dealing with a single leak, but with a model that is constantly changing its knowledge while simultaneously trying to keep track of where and when it learned what.

Lalam: This shows us that our commitment to privacy has to be more than just protecting data storage; it must protect the very representation of knowledge itself, ensuring that the way we express what a model knows is also trustworthy.

The solutions and their trade-offs: Tom: The paper introduces two clever ways to fix this leakage, which they call O learned and O prior. These are the core of the proposed fixes for our audience today.

Jane: One is using a separate DP mechanism to learn the labels themselves, which is called O learned. It’s like running an extra layer of noise just on the classification choices before we ever show them to a third party.

Meng: If we use this method, it means that even if a class exists in our private data, its presence in the output is noisy and unpredictable, which prevents anyone from confirming the existence of that class based on the labels alone.

Lu: I see a very powerful distinction here because we are not just splitting our total privacy budget for training weights; we’ are partitioning it between two distinct protection layers: selecting the labels and training them.

Lalam: This suggests a future where AI models aren't just trained to be accurate, but to present their knowledge in a way that is inherently safe and reliable, aligning with global values of trust.

Tom: The second method, O prior, uses public prior knowledge—a big set of labels known ahead of time—to decouple the output from the sensitive data.

Jane: So, instead of letting the labels come directly from a specific task's data, we use that pre-existing global list and re-map or drop any conflicting private labels we find in the training set.

Meng: Using O prior is a very practical way to keep things stable because it doesn't require us to find an unknown label in a specific task; it relies on known, pre-approved knowledge.

Lu: This approach allows the model to handle continuous streams of data even if we don't know all the possible classes ahead of time, which is crucial for future scalability.

Conclusion and final thoughts: Tom: So, wrapping up our discussion on "Privacy Leakage via Output Label Space and Differentially Private Continual Learning," we’ve seen that the output label space can be a massive privacy risk in sequential AI learning.

Jane: It's reassuring to know that the authors aren't just pointing out problems; they are providing viable solutions like O learned and O prior for managing this inherent risk.

Meng: I’m curious about the trade-offs, though; since both methods require either splitting our privacy budget or managing a large label space, how does that impact real-world efficiency? The data in the paper is really helpful here.

Lu: The authors provide guidance on which method to use based on whether you have prior knowledge or if your classes are very large, helping us choose the optimal path for complex environments.

Lalam: I hope that when we look at this research, we see it not as a failure of AI design, but as an opportunity to create a more robust future where privacy and trust are built in.

Tom: It seems like if classes are massive and we have no prior knowledge, Release Labels is the better choice for us.

Jane: And Public Labels are solid when you already have that broad set of known labels available for use in a scalable way.

Meng: We must keep these strategies in mind because they offer concrete ways to deploy secure, continuous AI systems that handle the data stream correctly.

Lu: We need systems that are both robust against forgetting and robust against this specific type of leakage, and the solutions provided by this paper are a big part of that complex engineering task.

Lalam: This work is a critical step toward ensuring that trust in AI isn' isn't just an aspiration but a verifiable reality through "Privacy Leakage via Output Label Space and Differentially Private Continual Learning." It fosters a culture where transparency is paramount.

Authors not visible in provided pages.

cs.LG, cs.CR

Submitted: 2026-08-24

Updated: 2026-08-25

Comments: 53 pages, 16 figures. Published in Transactions on Machine Learning Research (08/2026)

Code: https://github.com/PROBIC/private-continual-learning

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: the information leakage inherent in the output label space during continual learning (CL) tasks.

Key concepts

Privacy Leakage via Output Label Space
This is a dangerous vulnerability where an attacker uses prior knowledge of small differences between training datasets to accurately guess which dataset was used. The leakage occurs not just in the data itself but in how these subtle differences manifest in the classifier's output labels.
Continual Learning (CL)
CL describes a scenario where a model is learning from continuous streams of tasks. The challenge involves managing 'catastrophic forgetting' while ensuring the model constantly updates its knowledge base without leaking sensitive information across the entire sequence of tasks.
$O_{learned}$
This solution uses a separate Differential Privacy (DP) mechanism applied to the labels themselves. It adds noise to classification choices before they are shown to outsiders, preventing anyone from confirming the existence of any specific class based solely on the output.
$O_{prior}$
This method utilizes a large set of public prior knowledge, or a pre-existing list of labels. Instead of using labels directly from sensitive data, it re-maps or drops any conflicting private labels found in the training set using this known global list.

Terminology

Summary

Overview

The paper, titled Privacy Leakage via Output Label Space and Differentially Private Continual Learning, addresses a critical yet overlooked vulnerability in privacy-preserving machine learning: the information leakage inherent in the output label space during continual learning (CL) tasks. While Differential Privacy (DP) is commonly applied to model weights to protect training data, the authors demonstrate that the very structure of the model's output layer—specifically how new labels are introduced as a model learns new tasks—acts as a potent privacy side-channel. This leakage can lead to catastrophic privacy failures, even when standard DP algorithms are strictly implemented during weight updates.

The Core Problem: Label Space as a Side-Channel

In Continual Learning, models must adapt to new tasks sequentially. A common practice is to expand the model's output layer to accommodate new classes introduced in subsequent tasks. The researchers identify that naively adding the new labels to the model’s output label space during training can lead to a catastrophic privacy failure.

Essentially, even if the weights are protected by DP, the mere presence or modification of specific labels in the output space provides an adversary with high-confidence information about the sensitive categories present in the training data. This side-channel allows for membership inference or attribute leakage that bypasses traditional weight-based DP protections.

Proposed Mitigation Strategies

To counter this vulnerability, the authors propose two primary defense mechanisms designed to decouple the sensitive label information from the model's output structure:

  1. Release Labels (DP Label Release): This method involves applying an optimal Differential Privacy mechanism to the labels themselves before they are integrated into the model's output space. By adding controlled noise to the label identities, the authors ensure that an observer cannot definitively determine which specific sensitive classes were present in a given batch or task.

  2. Public Labels: This strategy utilizes a large, pre-defined public label space. Instead of expanding the output layer with new, sensitive labels for every task, the model maps tasks to a broader set of non-sensitive or public labels, thereby masking the introduction of private class information through structural changes in the output layer.

Methodology and Experimental Setup

The researchers evaluated their proposed methods by adapting two distinct families of Continual Learning algorithms to ensure the findings were not limited to a specific architecture:

  • Prototype Classifiers (Cosine Similarity Classifier): A method that uses feature embeddings compared against class prototypes via cosine similarity.

  • Expandable Parameter-Efficient Fine-Tuning (PEFT) Adapters: An approach that adds small, trainable modules to a pre-trained model, allowing for efficient task adaptation.

The empirical evaluation was conducted on two challenging benchmarks: Split-CIFAR-100 and Split-ImageNet-R. The performance was measured across various privacy budgets (epsilon) and compared against several baselines, including a Naive Baseline (no label protection) and a Full Data Baseline.

Key Findings and Recommendations

The study provides nuanced insights into the trade-offs between utility (accuracy), memory, and privacy:

  • Performance vs. Privacy: The proposed models consistently achieve higher accuracy under DP constraints compared to previous state-of-the-art methods, while simultaneously maintaining a stronger privacy model by mitigating the label side-channel.

  • Method Selection based on Context:

  • Release Labels is recommended when the number of classes is large or when there is no prior information available regarding the label distribution.

  • Public Labels serves as a robust alternative in controlled environments.

  • Classifier-Specific Insights:

  • The PEFT Ensemble tends to outperform the Cosine Classifier in non-blurry settings (where task boundaries and features are distinct).

  • The Cosine Classifier is the preferred choice when privacy requirements are extremely strict, computational/memory resources are limited, and the pre-training data distribution closely aligns with the sensitive dataset.

  • Theoretical Contribution: Beyond empirical results, the authors introduce a formalization of task-wise DP, providing a rigorous mathematical framework to reason about privacy guarantees in the specific context of sequential task learning.

Conclusion

The paper concludes that protecting model weights is insufficient in continual learning scenarios where the output space evolves. By addressing the label space side-channel through DP label release or public label mapping, practitioners can achieve a much more robust privacy guarantee without sacrificing significant model utility.

Improvements for AI systems

To improve AI systems based on this research, I would implement the following architectural upgrades:

  1. Implement a dual-budget privacy mechanism for Continual Learning (CL) systems that explicitly allocates a portion of the differential privacy (DP) budget to a Label Release Mechanism (using noisy thresholding via the private partition selection method). This prevents label space leakage, where an attacker can infer if a specific class was present in a training batch simply by observing the model's output layer.

  2. Integrate Data-Independent Prior Label Sets into classification heads for streaming data applications. Instead of updating the label space based on observed sensitive data, the system will use a large, public, fixed set of potential labels and map new task labels to this prior set using a remapping function (e.g., hierarchical grouping).

  3. Deploy PEFT-Ensemble architectures for privacy-preserving lifelong learning. Instead of fine-tuning an entire model—which is computationally expensive and prone to catastrophic forgetting under DP noise—the system will utilize Parameter-Efficient Fine-Tuning (specifically FiLM adapters) combined with an ensemble of task-specific heads.

By implementing these specific improvements, the resulting AI system will be able to:

  1. Perform continuous learning on sensitive, evolving data streams (e.g., medical diagnostics or financial transactions) without compromising individual privacy through output side-channels.

  2. Maintain high predictive accuracy and low forgetting rates even when operating under strict privacy constraints (low epsilon values).

  3. Scale to massive, multi-task environments with minimal memory overhead by only storing lightweight task-specific adapters rather than full model checkpoints for every historical task.

Abstract

Differential privacy (DP) is a formal privacy framework that enables training machine learning (ML) models while protecting individuals' data. As pointed out by prior work, ML models are part of larger systems, which can lead to so-called privacy side-channels even if the model training itself is DP. We identify the output label space of a classification model as such a privacy side-channel and show a concrete privacy attack that exploits it. The side-channel becomes highly relevant in continual learning (CL), where the output label space changes over time. To reason about privacy guarantees in CL, we introduce a formalisation of DP for CL, which also clarifies how our approach differs from existing approaches. We propose and evaluate two methods for eliminating this side-channel: applying an optimal DP mechanism to release the labels in the sensitive data, and using a large public label space. We explore the trade-offs of these methods through adapting pre-trained models. We demonstrate empirically that our models consistently achieve higher accuracy under DP than previous work over both Split-CIFAR-100 and Split-ImageNet-R, with a stronger privacy model.

Related papers