Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models".
Nadia: Fine-tuning Large Language Models (LLMs) can reawaken latent privacy risks, allowing previously learned private associations to become substantially more recoverable even without genuine private supervision.
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: So, we're diving into "Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models," and I'm curious what this means for the people listening who use these models every day.
Elias: It's a paper by Li et al., and the authors are looking at how fine-tuning can unexpectedly bring back private information that was supposed to be locked away.
Nadia: Exactly, and they suggest that this recovery doesn't always need the original private training data anymore, which is a significant point because that data is often unavailable to an attacker.
Priya: From a privacy measurement standpoint, I'm interested in how much of this recovered information we are actually talking about—is it just text or are we looking at deeper associations?
Nadia: That’s a fair question, Priya; the core idea is that previous attacks needed genuine private supervision drawn from the original training set, but this paper argues you can recover those same associations using only what the model generates.
Elias: And they propose this technique called ReGap as a way to do it by creating candidate answers and then picking the best ones based on how likely those tokens are under the frozen model before making any further adaptation.
Priya: So, if I'm understanding correctly, this isn't just about finding old text, but about recovering hidden patterns or associations that the model learned during its fine-tuning process?
Nadia: That’s right; they are focusing on entity-attribute associations that might be directly relevant to privacy, which goes beyond just verbatim sequences.
Elias: And they found this works across six models—GPT-two OPT, and Qwen3—which gives us a pretty broad look at the consistency of this phenomenon.
Priya: That breadth is important; if it holds up across different model architectures, it suggests the mechanism isn't tied to one specific type of model design.
Nadia: It really shows that post-training updates can create a new attack surface, and this paper lays out how an attacker can exploit that without having access to the original source material.
Elias: The implication is that we need to seriously re-evaluate post-deployment security because we might not be as safe as we think after a simple fine-tuning session.
The paper's summary: Nadia: Moving on from the authors, the paper summarizes this idea by stating that previous research focused on whether later updates could restore access to information that was hard to get directly, but they introduce a new attacker capability.
Elias: They argue that while studies have shown fine-tuning can reverse apparent forgetting after machine unlearning, this work looks at how malicious model-side interventions can increase privacy leakage during subsequent fine-tuning.
Nadia: Specifically, they contend that previous research often assumes the attacker has access to authentic private records from the target training distribution, which is not true in this scenario.
Priya: So, the summary is emphasizing that we need a method to recover associations learned *before* post-training without needing those original private records as supervision for the attack itself?
Elias: Precisely; they point out that memorization extends beyond just sequences to entity-attribute associations that are relevant to privacy, which is a key distinction.
Nadia: And they suggest using LLM-generated candidates as sufficient supervision because task structure and answer token likelihood can guide the recovery process.
Priya: What does this imply for privacy research? Does it mean we need to look at structural patterns within the model's outputs rather than just looking at the raw training data distribution?
Elias: They propose ReGap as a specific method: generating candidates, selecting them by negative log-likelihood under the frozen target model, and then adapting via low-rank adaptation.
Nadia: It’s a data-free attack, meaning it requires neither target answers nor auxiliary genuine private supervision during the candidate generation or selection phases.
Priya: That's a big deal for practical privacy auditing; if we can recover information this way, then simply releasing a model isn't the end of privacy protection.
The paper's improvements: Nadia: Now, regarding the specific mechanism they propose, the authors detail the ReGap pipeline as a three-stage process: Generate–Select–Adapt.
Elias: First, candidates are constructed by creating task-consistent queries using auxiliary keys and then querying the frozen target model to get one greedy answer and five stochastically sampled completions per query.
Priya: And what about that filtering stage? How do they know which of those generated answers are actually valid responses we should keep for the next step?
Nadia: They filter these outputs based solely on output validity and format, yielding a pool of distinct query-answer candidates. That helps narrow down the search space significantly.
Elias: The second stage is selection, where they compute a frozen support score using the negative log-likelihood under the frozen target model for each candidate pair.
Priya: So they are essentially ranking potential private associations based on how well they fit what the existing model already knows about correct answers?
Nadia: Exactly; they keep the top K candidates with the smallest selection loss to form their adaptation dataset, Daux. This is where the model learns from itself in a targeted way.
Elias: Finally, a fresh LoRA adapter is trained on this selected supervision using only causal cross-entropy focused just on those generated answer tokens.
Nadia: That’s the improvement: they are using model-generated supervision to train a fresh LoRA adapter, achieving recovery gains of six to twenty-one percentage points over the post-training target model across six models.
Conclusion: Nadia: To wrap up, the paper demonstrates that fine-tuning can indeed restore access to previously learned private associations without requiring genuine private supervision from the original training dataset.
Elias: The key finding is that by leveraging task structure and answer token likelihoods through ReGap, an attacker can achieve recovery gains of six to twenty-one percentage points on models like GPT-two OPT, and Qwen3.
Priya: And what this means for the data we're looking at is that the recovery persists even when adaptation identities are disjoint from all memorized or evaluation identities, provided there are no exact target answers in the supervision used.
Nadia: It also showed that prior exposure can increase recovery from forty-two point zero percent to sixty-three point zero percent on a previously exposed checkpoint, but this gain disappears when you use a matched target-unexposed control where the recovery stays at thirteen point zero percent.
Elias: The study does suggest that candidate selection acts as an amplifier rather than being the sole source of recovery, with frozen-model selection providing an additional gain based on the model's own structure.
Priya: I think this points to a need for privacy evaluation methods that assess both what a released model reveals directly and whether subsequent customization can make previously learned information accessible again through these self-generated means.
Jianhong Li, Jiahao Chen, Yuwen Pu, Chunyi Zhou, Oubo Ma, Zhou Feng, Hangtao Zhang, Jichao Bi
Chongqing University
cs.CR
Submitted: 2026-10-01
Updated: 2026-10-01
Comments: 34 pages
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 92/100
The gist: Fine-tuning Large Language Models (LLMs) can reawaken latent privacy risks, allowing previously learned private associations to become substantially more recoverable even without genuine private
Key concepts
- ReGap
- A three-stage attack pipeline: Generate, Select, Adapt. It uses the target model to create potential answers (Generate), ranks them based on how likely they are to be correct according to the frozen model (Select), and then trains a small adapter using only these selected candidates (Adapt).
- Answer-token likelihood
- This is a score used during selection. It measures how probable a specific token or sequence of tokens is as an answer, calculated using the frozen target model. Candidates with the lowest loss (highest likelihood) are chosen for further training.
- Low-rank adaptation (LoRA)
- A lightweight fine-tuning technique used to train a small adapter on selected supervision data. This allows the attacker to customize the large language model without needing access to all original private records, enabling them to reintroduce learned associations.
Terminology
Summary
Fine-tuning Large Language Models (LLMs) can reawaken latent privacy risks, allowing previously learned private associations to become substantially more recoverable even without genuine private supervision. This research introduces ReGap, a data-free attack that recovers these associations by generating model-generated candidates and selecting them based on answer-token likelihood under the frozen target model before performing low-rank adaptation.
The Gist
ReGap recovers previously learned private associations by generating task-consistent candidates from the target model, ranking them by answer-token likelihood under the frozen model, and training a fresh LoRA adapter on the selected supervision, achieving recovery gains of 6–21 percentage points over the post-training target model across six models.
The Threat Model and Objective
The paper investigates whether previously learned associations can be recovered from a released post-training model without access to original private records. The attacker knows the association schema and answer format, can query the target model, evaluate token-level likelihoods, and fine-tune it. The objective is to determine if an attacker can achieve a recovery gain of ∆rec = Rec(Ma) − Rec(Mh) > 0 using only model-generated supervision without externally supplied private records. The attack requires neither target answers nor auxiliary genuine private supervision during candidate generation, selection, or adaptation.
The ReGap Pipeline
ReGap employs a three-stage pipeline: Generate–Select–Adapt. First, candidates are generated by constructing task-consistent queries using auxiliary keys and querying the frozen target model to generate one greedy answer and five stochastically sampled completions per query. These outputs are filtered for invalid responses or malformed answers based only on output validity and format, yielding a pool of distinct query–answer candidates. Second, selection is performed by computing the frozen support score, which is defined as the negative log-likelihood under the frozen target model: lz(0) = −1/y Σ X y t=1 log pMh(yt q, y<t). The top K candidates with the smallest selection loss are retained to form the adaptation dataset, Daux. Finally, a fresh LoRA adapter is trained on this selected supervision using causal cross-entropy only on the generated answer tokens.
Key Findings and Controls
The research systematically characterizes recovery across six models (GPT-2, OPT, Qwen3) and various controls. One round of ReGap improves target-association recovery by 6–21 percentage points over the target model. This recovery persists when adaptation identities are disjoint from all memorized and evaluation identities, with no exact target answers appearing in the generated or selected supervision. A matched prior-exposure control demonstrates that the same trained adapters increase recovery from 42.0% to 63.0% on a previously exposed checkpoint, but produce no gain on a matched target-unexposed control (13.0%).
Structural Boundaries and Limitations
The study establishes boundaries for the effectiveness of model-generated supervision. Recovery weakens sharply when the underlying structure is removed; specifically, in an opaque random-binding control where lexical relationships are removed between names and answers, recovery drops from 95.0% to 1.0% after non-target continual training and reaches only 2.0% after ReGap-Abs, while the matched target-unexposed branch remains at 0%. Furthermore, explicit Negative Preference Optimization (NPO) unlearning yields almost no additional recovery after ReGap in stress tests. The effectiveness of model-generated supervision is shown to be distribution dependent, as removing dominant corporate-domain regularity weakens the effect on non-Enron domains. The analysis also suggests that candidate selection acts as an amplifier rather than the sole source of recovery, with frozen-model selection providing an additional, model-dependent gain.
Conclusion
The findings demonstrate that fine-tuning can restore access to previously learned private associations without genuine private supervision. Recovery depends on the associations being recovered and the updates the model has undergone; prior exposure affects post-adaptation recoverability under lightweight adaptation, while recovery weakens sharply for opaque random bindings and after explicit NPO unlearning. Privacy evaluation should assess both what a released model reveals directly and whether subsequent customization can make previously learned information accessible again, including when the customization data come from the model itself.
AI Use Statement
Generative AI tools were used during this work to assist with research discussion, experimental design feedback, code development and debugging, result interpretation, translation, and manuscript drafting and editing. They were also used to help organize experimental records and improve the clarity and structure of the presentation. All AI-assisted outputs were manually reviewed by the authors. Experimental results, numerical claims, methodological descriptions, and citations were checked against the corresponding code, logs, source files, and original references before inclusion in the manuscript. The authors take full responsibility for the final content of the paper.
References
Atilla Akkus et al., "Generated data with fake privacy:
Improvements for AI systems
Based on the scientific paper, here are specific improvements that can be made to existing AI systems:
-
Improve Privacy Resilience in Fine-Tuning Pipelines:
-
Enhance Post-Adaptation Privacy Auditing Capabilities:
-
Develop Data-Free Recovery Mechanisms for Latent Information:
-
Implement Adaptive Supervision Strategies for Model Customization:
Specific capabilities these improvements enable:
-
The system can be more resilient to privacy leakage introduced during routine model fine-tuning (e.g., domain adaptation, instruction tuning). It can actively detect and mitigate the reawakening of previously memorized private associations that become accessible after post-training updates, even when the fine-tuning data is non-target or model-generated.
-
The system can perform
privacy audits
on a released model checkpoint to determine not just direct leakage, but also its susceptibility to future recovery attacks via adaptation. This allows for proactive risk assessment across the entire model lifecycle (pre-deployment through subsequent fine-tuning). -
The system can recover previously learned private associations using only the target model's own generated outputs as supervision, without needing access to original private training data or exact target answers. This enables the recovery of hidden knowledge based purely on structural patterns and answer likelihoods present in the model's current state.
-
The system can utilize a structured
Generate-Select-Adapt
pipeline where it generates candidate examples from itself, selects them based on their likelihood to be correct (using frozen target model scores), and then adapts the model using only that selected, self-generated supervision. This allows for targeted knowledge recovery with minimal or no access to sensitive ground truth labels.
Sources
- Quantifying Memorization Across Neural Language Models
- Profiling for Pennies: Unveiling the Privacy Iceberg of LLM Agents
- LoRA: Low-Rank Adaptation of Large Language Models
- Measuring Forgetting of Memorized Training Examples
- Towards Continual Knowledge Learning of Language Models
- LLM-PBE: Assessing Data Privacy in Large Language Models
- Pointer Sentinel Mixture Models
- Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
- Qwen3 Technical Report
- OPT: Open Pre-trained Transformer Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs