One Risk Down, Another Up: Cross-Risk Interactions Induced by LLM Defenses
cs.CR
Submitted: 2025-10-09
Updated: 2026-09-01
Comments: Under Review
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Language Models are Few-Shot Learners
- Knowledge Neurons in Pretrained Transformers
- GPT-4 Technical Report
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
- Toxicity in ChatGPT: Analyzing Persona-assigned Language Models
- Queer People are People First: Deconstructing Sexual Identity Stereotypes in Large Language Models
- Disclosure and Mitigation of Gender Bias in LLMs
- Dialect prejudice predicts AI decisions about people's character, employability, and criminality
- The Llama 3 Herd of Models
- Are Large Pre-Trained Language Models Leaking Your Personal Information?
- Toy Models of Superposition
- No Free Lunch with Guardrails
- A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
- Multi-step Jailbreaking Privacy Attacks on ChatGPT
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- LLM-PBE: Assessing Data Privacy in Large Language Models
- Debiasing Algorithm through Model Adaptation
- DeepSeek-V3 Technical Report
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs