Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
cs.LG, cs.AI
Submitted: 2026-07-31
Updated: 2026-09-26
Terminology
Sources
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- Defending Against Unforeseen Failure Modes with Latent Adversarial Training
- Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
- MixAT: Combining Continuous and Discrete Adversarial Training for LLMs
- Toy Models of Superposition
- The Llama 3 Herd of Models
- Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
- Impact of Positional Encoding: Clean and Adversarial Rademacher Complexity for Transformers under In-Context Regression
- Scaling Trends in Language Model Robustness
- Scaling Laws for Neural Language Models
- A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
- Towards Deep Learning Models Resistant to Adversarial Attacks
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Pruning Convolutional Neural Networks for Resource Efficient Inference
- Softmax is $1/2$-Lipschitz: A tight bound across all $\ell_p$ norms
- Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
- Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
- Open Problems in Mechanistic Interpretability
- Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks