High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection
summary
The gist
Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples, and this vulnerability can persist even when quality-based data selection is
In short
This research investigated whether selecting high-quality data removes poisoning samples from LLMs during fine-tuning. The study found that while selection removes overtly harmful examples, retaining high-quality samples can still degrade safety alignment because they possess training patterns similar to harmful data, causing safety conflicts to accumulate.
Key concepts
- Quality-based Data Selection
- This method filters a dataset by assigning quality scores to samples and keeping only those above a certain threshold. The researchers found this successfully removes many clearly harmful examples, but it doesn't guarantee the remaining high-quality data is safe.
- Harmful-Gradient Influence Proxy P(x)
- This metric measures how closely a specific data sample's gradient update resembles the direction of known harmful samples. A higher score indicates that the sample's training effect is pushing the model toward a harmful outcome, even if it has a good quality score.
- Cumulative Safety Drift
- This describes how safety issues worsen over time during fine-tuning. The paper shows that if the accumulated negative updates from retained samples are strong enough, they cause the model's safety objective to degrade predictably, leading to reduced safety performance.
Terminology used across episodes
This episode discusses
- High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection · Paper Radio
- The Llama 3 Herd of Models · Paper Radio
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- WanJuan-CC: A Safe and High-Quality Open-sourced English Webtext Dataset
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen3 Technical Report
- TextGrad: Automatic "Differentiation" via Text
- Understanding and Preserving Safety in Fine-Tuned LLMs
- A Survey of Large Language Models
The paper
High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection · Read on arXiv
Kaiyang Li, Jiahao Chen
School of Big Data & Software Engineering, Chongqing University · College of Computer Science and Technology, Zhejiang University
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection".
Nadia: Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples, and this vulnerability can persist even when quality-based data selection is employed.
Elias: First, who's behind it and why it matters.
Paper summary: Elias: I agree, Nadia; the authors systematically evaluate both the filtering effects against poisoning and the downstream safety impact of those retained data points. They reveal that while selection methods remove many malicious examples by assigning low quality scores to them, some retained high-quality samples still carry a harmful-like training pattern at the layerwise gradient level.
Priya: It’s important to understand why this matters for privacy and measurement research: it shows that the safety issue isn't just about getting bad data in; it's about how benign-looking data can subtly introduce safety conflicts during the actual fine-tuning phase. This suggests our metrics need to capture these internal gradient relationships, not just the initial quality score.
Nadia: Precisely, Priya. The paper highlights a practical vulnerability where safety-degrading influence can pass through quality-based selection via those retained high-quality samples that are still active in the model updates. This is a key finding because it means our current defenses against poisoning might be incomplete if they don't account for this downstream effect on alignment.
Elias: The central question they address is whether high-quality data truly translates to safety during fine-tuning, and their analysis shows that retained high-quality data can still degrade safety alignment, evidenced by samples exhibiting predominantly negative gradient similarities that oppose the intended safety direction.
Priya: That concept of "safety-conflicting updates" sounds like a huge problem for privacy researchers too, because it means the model's learned behavior is being subtly steered away from its intended safe state by these seemingly good training examples. What kind of behavioral drift are we talking about?
Nadia: We’re talking about a subtle erosion of safety alignment where the model starts exhibiting unsafe responses even though it was trained on what looked like high-quality, benign material. The paper investigates the layerwise gradient relationships between training samples and both harmful and safe anchors to find the source of this drift.
Elias: By analyzing those patterns in a low-dimensional space, they found that some high-quality samples display gradient patterns similar to those of explicitly harmful data, particularly in the lower and middle layers. This offers a parameter-space perspective on why this happens during fine-tuning.
Priya: That layer specificity is telling; it suggests that the safety degradation isn't uniform across the model but is localized within specific parts of the network architecture, which has implications for targeted mitigation strategies.
Nadia: Exactly, Priya. They are suggesting that these specific samples might induce harmful-like updates precisely in those lower and middle layers, which explains why they weaken overall safety alignment during fine-tuning. This provides a much clearer picture than just looking at the aggregate data quality score.
Elias: Taken together, the findings of "High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection" expose a practical vulnerability: even when quality selection removes many malicious samples, it doesn't prevent high-quality data from serving as carriers of safety-degrading influence during downstream fine-tuning.
Priya: It’s a strong warning that we can't rely solely on initial data curation to ensure final model safety; the mechanism of influence needs to be understood throughout the training lifecycle.
Conclusion: Elias: I think their work by Kaiyang Li and colleagues forces us to realize that we need to look past surface-level metrics like quality scores when evaluating model safety; the actual gradient dynamics during fine-tuning are what reveal these hidden risks, even when the input data seems perfectly fine.
Priya: From a broader impact view, this suggests that developing robust safety protocols for large language models needs to integrate analysis of how data influences internal model gradients, not just checking the raw inputs before training starts. If this is true, it fundamentally changes how we approach adversarial robustness in these systems.
Nadia: That's right; the implication is that future research and development need to focus on understanding those specific gradient patterns—like those in the lower and middle layers they found—to build better safety mechanisms that are sensitive to this type of subtle, persistent influence.
Elias: It’s a call for a more holistic approach where we don't just filter out bad data but actively monitor how retained high-quality data interacts with the model's safety objectives as it learns. This paper gives us tools to examine that interaction at a granular level.
Priya: So, ultimately, the paper emphasizes that quality selection is necessary but not sufficient for ensuring model safety; we have to look at the actual training dynamics to see if those retained samples are actually introducing harmful-like updates into the model's learned behavior.
Nadia: That’s exactly what it means—the vulnerability lies in assuming that a high quality score equals a safe update, which this paper shows is not the case when looking at layerwise gradients. We need to be much more skeptical of simple data quality metrics when assessing downstream safety risks.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel