High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection".
Nadia: Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples, and this vulnerability can persist even when quality-based data selection is employed.
Elias: First, who's behind it and why it matters.
Paper summary: Elias: I agree, Nadia; the authors systematically evaluate both the filtering effects against poisoning and the downstream safety impact of those retained data points. They reveal that while selection methods remove many malicious examples by assigning low quality scores to them, some retained high-quality samples still carry a harmful-like training pattern at the layerwise gradient level.
Priya: It’s important to understand why this matters for privacy and measurement research: it shows that the safety issue isn't just about getting bad data in; it's about how benign-looking data can subtly introduce safety conflicts during the actual fine-tuning phase. This suggests our metrics need to capture these internal gradient relationships, not just the initial quality score.
Nadia: Precisely, Priya. The paper highlights a practical vulnerability where safety-degrading influence can pass through quality-based selection via those retained high-quality samples that are still active in the model updates. This is a key finding because it means our current defenses against poisoning might be incomplete if they don't account for this downstream effect on alignment.
Elias: The central question they address is whether high-quality data truly translates to safety during fine-tuning, and their analysis shows that retained high-quality data can still degrade safety alignment, evidenced by samples exhibiting predominantly negative gradient similarities that oppose the intended safety direction.
Priya: That concept of "safety-conflicting updates" sounds like a huge problem for privacy researchers too, because it means the model's learned behavior is being subtly steered away from its intended safe state by these seemingly good training examples. What kind of behavioral drift are we talking about?
Nadia: We’re talking about a subtle erosion of safety alignment where the model starts exhibiting unsafe responses even though it was trained on what looked like high-quality, benign material. The paper investigates the layerwise gradient relationships between training samples and both harmful and safe anchors to find the source of this drift.
Elias: By analyzing those patterns in a low-dimensional space, they found that some high-quality samples display gradient patterns similar to those of explicitly harmful data, particularly in the lower and middle layers. This offers a parameter-space perspective on why this happens during fine-tuning.
Priya: That layer specificity is telling; it suggests that the safety degradation isn't uniform across the model but is localized within specific parts of the network architecture, which has implications for targeted mitigation strategies.
Nadia: Exactly, Priya. They are suggesting that these specific samples might induce harmful-like updates precisely in those lower and middle layers, which explains why they weaken overall safety alignment during fine-tuning. This provides a much clearer picture than just looking at the aggregate data quality score.
Elias: Taken together, the findings of "High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection" expose a practical vulnerability: even when quality selection removes many malicious samples, it doesn't prevent high-quality data from serving as carriers of safety-degrading influence during downstream fine-tuning.
Priya: It’s a strong warning that we can't rely solely on initial data curation to ensure final model safety; the mechanism of influence needs to be understood throughout the training lifecycle.
Conclusion: Elias: I think their work by Kaiyang Li and colleagues forces us to realize that we need to look past surface-level metrics like quality scores when evaluating model safety; the actual gradient dynamics during fine-tuning are what reveal these hidden risks, even when the input data seems perfectly fine.
Priya: From a broader impact view, this suggests that developing robust safety protocols for large language models needs to integrate analysis of how data influences internal model gradients, not just checking the raw inputs before training starts. If this is true, it fundamentally changes how we approach adversarial robustness in these systems.
Nadia: That's right; the implication is that future research and development need to focus on understanding those specific gradient patterns—like those in the lower and middle layers they found—to build better safety mechanisms that are sensitive to this type of subtle, persistent influence.
Elias: It’s a call for a more holistic approach where we don't just filter out bad data but actively monitor how retained high-quality data interacts with the model's safety objectives as it learns. This paper gives us tools to examine that interaction at a granular level.
Priya: So, ultimately, the paper emphasizes that quality selection is necessary but not sufficient for ensuring model safety; we have to look at the actual training dynamics to see if those retained samples are actually introducing harmful-like updates into the model's learned behavior.
Nadia: That’s exactly what it means—the vulnerability lies in assuming that a high quality score equals a safe update, which this paper shows is not the case when looking at layerwise gradients. We need to be much more skeptical of simple data quality metrics when assessing downstream safety risks.
Kaiyang Li, Jiahao Chen
School of Big Data & Software Engineering, Chongqing University · College of Computer Science and Technology, Zhejiang University
cs.CR
Submitted: 2026-10-01
Updated: 2026-10-01
Comments: 15 pages, in submission
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 92/100
The gist: Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples, and this vulnerability can persist even when quality-based data selection is
Key concepts
- Quality-based Data Selection
- This method filters a dataset by assigning quality scores to samples and keeping only those above a certain threshold. The researchers found this successfully removes many clearly harmful examples, but it doesn't guarantee the remaining high-quality data is safe.
- Harmful-Gradient Influence Proxy P(x)
- This metric measures how closely a specific data sample's gradient update resembles the direction of known harmful samples. A higher score indicates that the sample's training effect is pushing the model toward a harmful outcome, even if it has a good quality score.
- Cumulative Safety Drift
- This describes how safety issues worsen over time during fine-tuning. The paper shows that if the accumulated negative updates from retained samples are strong enough, they cause the model's safety objective to degrade predictably, leading to reduced safety performance.
Terminology
Summary
Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples, and this vulnerability can persist even when quality-based data selection is employed. The core finding reveals that while selection removes many overtly harmful samples, retained high-quality samples can still degrade model safety alignment due to their harmful-like training patterns at the layerwise gradient level.
Research Questions
The study addresses three primary research questions concerning the interaction between quality-based data selection and fine-tuning poisoning:
-
Can quality-based data selection exclude poisoning samples from the final fine-tuning dataset? The researchers found that these methods "remove some poisoning examples, particularly those containing explicitly harmful content or receiving low quality scores, thereby reducing the amount of malicious data in the final fine-tuning corpus and mitigating poisoning.
However, it remains unclear whether
the high-quality samples that survive selection can still weaken the model’s safety alignment during fine-tuning." -
Do high-quality samples retained after selection truly mean safety during fine-tuning? Analysis showed that
retained high quality data can still degrade safety alignment,
evidenced by retained samples exhibitingpredominantly negative gradient similarities, indicating safety-conflicting updates that oppose the safety-alignment direction and may accumulate during fine-tuning.
-
Why can some high-quality data still undermine safety alignment? The investigation into layer-wise gradients revealed that
some high-quality samples exhibit gradient patterns similar to those of explicitly harmful data, particularly in the lower and middle layers,
suggesting thatsuch samples may induce harmfullike updates and weaken safety alignment during fine-tuning.
Methodology: Bi-Stage Quality-Constrained Safety-Degradation Text Optimization (Bi-QSTO)
To systematically examine the exploitability of this threat, the authors propose Bi-QSTO, a bi-stage framework designed to optimize poisoned samples under an explicit quality constraint to survive selection while preserving their safety-degrading influence. This method evaluates each candidate sample using two signals: a harmful-gradient influence proxy signal P(x) from the target model and a relative quality signal Q(x) from the data selector.
Harmful-Gradient Influence Proxy:
The harmful-gradient influence proxy, P(x), is calculated as the cosine similarity between the candidate's gradient, g(x), and the mean gradient of a harmful anchor set, g¯h:
P(x) = g(x)⊤g¯h / (∥g(x)∥2 ∥g¯h∥2). A higher P(x) indicates that the candidate induces an update closer to the harmful anchor direction.
Relative Quality Signal:
The relative quality signal, Q(x), is defined as:
Q(x) = o(x) + (o(x) - 1/n Xj1 o(bj)), where o(·) denotes the overall quality rating. A higher Q(x) indicates both a favorable quality score and a larger advantage over the background data.
Bi-Stage Optimization Process
Bi-QSTO employs a two-stage optimization process to balance the conflicting objectives of maximizing harmful influence and satisfying quality constraints:
-
The first stage seeks stronger harmful-like influence by maximizing P(x):
Stage 1 seeks stronger harmful-like influence: x∗1 = arg max x∈X P(x).
This stage iteratively rewrites and evaluates sample variants to obtainstronger harmful-gradient influence.
-
The second stage refines the best variant from Stage 1 under the quality constraint while preserving its harmful gradient influence:
Stage 2 refines the best variant from Stage 1 under the quality constraint while preserving its harmfulgradient influence: x∗2 = arg max x∈X n P(x) − λ [max(0, τ − Q(x))]2 o.
The penalty term grows quadratically with the quality deficit if Q(x) is below a threshold τ, guiding the search toward higher-quality variants.
Key Findings and Theoretical Formalization
The theoretical analysis establishes a chain from selection survival to safety degradation through two key theorems:
Theorem 1 (Unavoidable Quality-Score Overlap):
This theorem proves that if poisoning preserves the quality score within a bounded distortion, then the two score distributions have overlapping support,
meaning that any threshold retained with a sufficient quality margin necessarily retains a nonzero fraction of poisoning samples.
This explains why observed overlap is not merely an artifact but a structural consequence of quality-preserving poisoning.
Theorem 2 (Cumulative Safety Drift):
This theorem characterizes the effect during fine-tuning, showing that safety-conflicting updates accumulate. It demonstrates that if the cumulative first-order safety-degrading influence exceeds the second-order curvature term, the safety-reference loss provably increases,
leading to degradation of the safety objective.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed this paper, High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection,
which introduces the Bi-Stage Quality-Constrained Safety-Degradation Text Optimization (Bi-QSTO) framework.
The core contribution is demonstrating that safety alignment in fine-tuned Large Language Models (LLMs) is vulnerable to poisoning even when quality filters are applied, because retained high-quality samples can induce harmful gradient patterns. The proposed solution, Bi-QSTO, optimizes poisoned samples to survive selection while preserving their safety-degrading influence.
Here are the specific improvements and what the improved AI system can achieve:
)
AI System Improvements Enabled by Bi-QSTO:
-
[] Improved Robustness Against Fine-Tuning Poisoning under Quality Filtering: The system will be significantly more resilient to poisoning attacks that evade explicit content moderation (i.e.,
benign-looking
malicious samples). -
[] Enhanced Safety Alignment Preservation During Downstream Training: The model will maintain its initial safety alignment (against harmful instructions) even after undergoing quality-based data selection for fine-tuning, as the Bi-QSTO method actively mitigates the safety degradation caused by retained high-quality poisoned data.
-
[] Adaptive Poisoning Generation: Instead of relying on static, overtly harmful samples, the system can generate
high-quality malicious
samples that are specifically engineered to survive quality filters while maximizing their potential to weaken safety alignment during fine-tuning (as shown by the success of Harmful-Seed optimization). -
[] Layer-Wise Gradient Awareness: The underlying mechanism utilizes layer-wise gradient analysis to understand how different data types (benign vs. malicious) affect the model's internal parameter space, allowing for a deeper understanding and targeted defense against safety drift during training.
)
What the Improved AI System Can Do (Specific Capabilities):
-
[] Fine-Tuning Jailbreak Resistance: The system can be fine-tuned on potentially compromised datasets (e.g., data sourced from the web or crowdsourced) without risking a
fine-tuning jailbreak
where benign but harmful samples subtly override safety guardrails during the alignment process. -
[] Defense Against Semantic Evasion Attacks: It can defend against sophisticated poisoning attacks that use semantically benign instructions or low-toxicity language, which previously bypassed simple content filters but are now identified and neutralized by Bi-QSTO optimization.
-
[] High-Fidelity Attack Deployment: Researchers can deploy highly effective adversarial training data that survives standard data cleaning pipelines, ensuring the model is exposed to the most challenging forms of poisoning before deployment.
-
[] Automated Data Selection Strategy Integration: The system can be integrated into real-world MLOps pipelines where quality scores are used for automated filtering, and Bi-QSTO can be used as a pre-processing step to
sanitize
the remaining high-quality samples before they enter the final fine-tuning set. -
[] Quantifiable Safety Drift Monitoring: By leveraging the gradient analysis insights, researchers can monitor whether retained high-quality data is causing measurable safety degradation (as quantified by the safety-reference loss in Theorem 2), enabling proactive intervention if necessary.
Sources
- The Llama 3 Herd of Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- WanJuan-CC: A Safe and High-Quality Open-sourced English Webtext Dataset
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen3 Technical Report
- TextGrad: Automatic "Differentiation" via Text
- Understanding and Preserving Safety in Fine-Tuned LLMs
- A Survey of Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs