GradSentry: Gradient Spectral Entropy for Backdoor Sample Filtering in Large Language Model Fine-Tuning
cs.CR
Submitted: 2026-05-26
Updated: 2026-08-31
Code: https://github.com/dongdongzhaoUP/GradSentry
Terminology
Sources
- GPT-4 Technical Report
- Lost in Backpropagation: The LM Head is a Gradient Bottleneck
- BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- ProtegoFed: Backdoor-Free Federated Instruction Tuning with Interspersed Poisoned Data
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs